Multi-Agent Systems in Production: Architecture Patterns That Hold Up
Short answer
A multi-agent system splits work between several specialised LLM agents coordinated by code or by an orchestrating agent. In production, the patterns that hold up are: a deterministic pipeline of agents, a router that sends tasks to specialists, an orchestrator that delegates to workers, and a generator–evaluator loop. Use them only when a single agent with good tools can't do the job, and engineer explicit state, limits, validation and tracing around every hand-off.
Key takeaways
- Start with a single agent; move to multiple agents only when the task clearly benefits.
- Prefer deterministic orchestration in code over agents freely chatting with each other.
- Persist state at every step so failures can resume instead of restarting.
- Set hard limits on steps, cost and time, and trace every hand-off.
Multi-agent systems are one of the most hyped ideas in AI engineering. Diagrams show a "CEO agent" delegating to a "research agent," a "writer agent" and a "critic agent," all collaborating like a small company.
Some of that works in production. A lot of it doesn't.
We've built and run multi-agent systems for financial analysis, document processing and operations automation. This guide covers the patterns that hold up, the engineering that makes them reliable, and the cases where you should avoid multiple agents altogether.
What is a multi-agent system?
A multi-agent system is an application where several AI agents — each with its own instructions, tools and responsibilities — work together on a task. Coordination comes from either:
- Deterministic code (a workflow or graph you define), or
- An orchestrator agent that decides which other agents to call and when.
Each agent is typically a language model call (or loop of calls) with a focused system prompt and a limited toolset. If you're new to agents generally, start with AI agents vs chatbots vs RPA.
When do you actually need more than one agent?
Our default advice: start with one agent and good tools. Move to multiple agents only when you hit one of these limits:
- Context overload. One agent needs so many instructions and tools that it starts choosing wrongly.
- Different capabilities. Subtasks need different models (a cheap fast one for classification, a strong one for analysis) or different permissions.
- Parallelism. Independent subtasks — researching ten suppliers — can run at the same time.
- Separation of duties. One agent generates, another checks, so mistakes are caught before they reach a user.
If none of these apply, a single agent will be simpler, cheaper and easier to debug.
Which multi-agent architecture patterns work in production?
| Pattern | How it works | Best for | Main risk |
|---|---|---|---|
| Sequential pipeline | Fixed chain: agent A → B → C | Document processing, multi-stage analysis | Errors compound across stages |
| Router | A classifier sends each task to one specialist | Support triage, mixed request types | Misrouting |
| Orchestrator–worker | An orchestrator plans and delegates subtasks, then combines results | Research, complex analysis, parallel work | Over-planning, cost blow-ups |
| Generator–evaluator | One agent produces, another critiques; loop until it passes | Content, code, extraction with strict rules | Endless loops |
| Hierarchical | Orchestrators manage sub-orchestrators | Very large workflows | Hard to debug; rarely needed |
Pattern 1: Sequential pipeline
The most reliable pattern, because the flow is fixed in code. Example: extract fields from a contract → classify each clause → compare against a playbook → summarise risks. Each agent has one job and a validated output schema. Our legal document intelligence write-up uses this approach.
Make it reliable: validate output between every stage so a bad extraction stops the pipeline rather than poisoning the next step.
Pattern 2: Router
A lightweight model classifies each request and routes it to a specialist agent with the right instructions and tools. This keeps each specialist's context small and focused.
Make it reliable: include an "unsure" route that goes to a human or a general agent, and evaluate the router separately — misrouting is a silent failure.
Pattern 3: Orchestrator–worker
An orchestrator agent breaks a goal into subtasks, delegates them to worker agents (often in parallel), then synthesises the results. Good for open-ended research and analysis where the steps can't be known in advance.
Make it reliable: cap the number of subtasks and total model calls; give workers narrow tools; require structured results so the orchestrator isn't summarising free text from free text.
Pattern 4: Generator–evaluator loop
One agent produces an output; another checks it against explicit criteria and either approves it or returns specific feedback. Effective for tasks with clear rules: data extraction, compliance checks, code generation.
Make it reliable: limit the loop to two or three rounds, and escalate to a human if it doesn't converge.
Should agents talk to each other freely?
Generally, no. Designs where agents converse in open-ended chat until they agree look impressive in demos but are hard to control: they loop, drift off-task, inflate costs and are difficult to debug.
In production, we prefer explicit hand-offs: each agent receives structured input and returns structured output, and code (or a tightly constrained orchestrator) decides what happens next. Frameworks like LangGraph model this as a graph of nodes and edges with explicit state, which makes the flow visible and testable.
How do you manage state in a multi-agent system?
State is where most multi-agent systems quietly break. Rules we follow:
- One source of truth. Shared state lives in a store (PostgreSQL, Redis), not scattered across agents' conversation histories.
- Structured state, not chat logs. Pass typed objects between agents:
{ invoice_id, extracted_fields, confidence, issues[] }. - Checkpoint every step. If step 4 fails, resume from step 4 instead of re-running (and re-paying for) steps 1–3.
- Pass only what's needed. Each agent gets the minimum context for its job, which improves accuracy and reduces cost.
- Idempotent actions. Anything that changes external systems must be safe to retry.
How do you handle failures across agents?
| Failure | Prevention |
|---|---|
| Agent returns malformed output | Schema validation with a retry that includes the validation error |
| Agent loops or never finishes | Hard limits on steps, tokens, time and cost per task |
| Error compounds across stages | Validation gates between stages; confidence thresholds |
| Tool call fails | Retries with backoff; fallback paths; clear error messages to the agent |
| Model provider outage or rate limit | Fallback to a second provider; queue with backpressure |
| Agent attempts a risky action | Least-privilege tools; human approval above thresholds |
How do you observe and debug multi-agent systems?
You can't fix what you can't see. Every production multi-agent system needs:
- End-to-end traces showing each agent, prompt, tool call, output, latency and cost, linked under one task ID
- Per-agent metrics, so you know which role is slow, expensive or error-prone
- Evaluation per agent and end-to-end, so a change to one agent doesn't silently break the whole
OpenTelemetry's generative AI conventions and LLM-specific tools such as Langfuse make this practical. See LLM observability in production and how to evaluate LLM applications.
How do agents share tools and data?
As the number of agents grows, so does the number of integrations. Rather than wiring tools into each agent, expose shared systems through a common layer — increasingly the Model Context Protocol (MCP) — with authentication and permissions enforced centrally. Each agent then gets access only to the tools its role needs.
How do you control cost in multi-agent systems?
- Use smaller, faster models for routing, classification and extraction; reserve frontier models for hard reasoning
- Cache shared instructions and context (prompt caching)
- Parallelise independent work to cut latency, but cap concurrency
- Set budgets per task and alert on outliers
More techniques in how to cut LLM API costs.
How we build multi-agent systems at Keyved
Our Fintech AI Suite and document-processing systems use multi-agent designs where they clearly help — parallel analysis, generator–evaluator checks, specialist routing — and single agents everywhere else.
All of them run on our platform foundation, which provides persistent state and checkpointing, retries and provider failover, and full tracing out of the box. That's what turns an impressive multi-agent diagram into a system that runs every day.
Considering a multi-agent design? See our AI agents and automation service, or talk to our engineers — we'll tell you whether you need several agents or one well-equipped one.
Frequently asked questions
What is a multi-agent system in AI?
A multi-agent system is an application where several AI agents, each with its own instructions, tools and responsibilities, work together on a task. They are coordinated either by deterministic code or by an orchestrator agent that delegates subtasks.
When should I use multiple agents instead of one?
Use multiple agents when a task has clearly separable subtasks that need different tools, instructions or models; when subtasks can run in parallel; or when a single agent's context becomes too large to manage reliably. Otherwise, a single agent is simpler, cheaper and easier to debug.
What frameworks are used to build multi-agent systems?
Common options include LangGraph, the OpenAI Agents SDK, the Claude Agent SDK, CrewAI, AutoGen and Microsoft's agent frameworks, as well as plain code with a workflow engine. The framework matters less than explicit state, limits, validation and observability.
Why do multi-agent systems fail?
Common causes are unclear responsibilities between agents, lost context at hand-offs, agents looping or arguing without limits, errors compounding across steps, and no tracing to see where things went wrong.
Are multi-agent systems expensive to run?
They usually cost more than a single agent because each agent makes its own model calls. Costs can be controlled by using smaller models for simpler roles, caching shared context and limiting the number of steps.