Skip to main content
Guides

How to design multi-agent systems

Reading time
12 minutes
Assumes
You've built a single-agent system and hit its limits
Updated
Sep 6, 2026

Start by not doing it

Multi-agent designs are appealing because they mirror how we organize people, and org charts feel like sound architecture. They usually aren't, for a system where the expensive resource is context rather than headcount.

Every boundary you introduce costs something concrete. Context gets lost at the handoff — the second agent knows what the first passed, not what it saw. Failures compound: three steps at 90% is 73% end to end. Debugging gets harder, because a wrong answer at the end could originate anywhere upstream. And latency and cost multiply with each hop.

A single agent with a clear instruction, good tools, and the relevant context outperforms a committee of specialists on most tasks people reach for multi-agent designs to solve. Exhaust that first — it is also much easier to evaluate.

The bar for a split

Split when a boundary already exists in the problem: a different actor, a different failure mode you need to isolate, a different approval requirement, a body of knowledge with its own retrieval rules, or a genuinely different task shape. Not when the prompt got long.

Boundaries that hold up

Five reasons to split that survive contact with production.

Different actors. Part of the work is done by a person. That is a real boundary and the handoff has to be explicit anyway.

Different approval requirements. One part reads and one part writes to the outside world. Splitting lets you gate the second without gating the first, and that alone can make an approval queue survivable.

Different failure isolation. You need one part to fail without taking the rest down, or you need to retry one part independently. Retrying a monolithic step that's already sent an email is not a retry.

Different knowledge to search. Policy documents, a product catalog and two years of support threads are three corpora with three retrieval strategies — different chunking, different queries, different rules about what counts as grounded. One agent holding all three searches all three on every question, and tuning retrieval for one makes it worse for the others. Isolating a corpus also keeps its results out of the caller's context: the sub-agent searches, and what comes back is a short grounded answer with citations rather than twenty passages the orchestrator has to read past.

Genuinely different shapes. Classification, extraction into a schema, and open-ended drafting are different tasks with different evaluation criteria. Separating them lets you measure each properly, which is usually the strongest practical argument.

The reason that doesn't hold up: "the prompt was getting long." A long prompt is a prompt problem. Splitting it distributes the confusion across two components and adds a handoff.

What a sub-agent actually is

Here is the part most architecture diagrams leave out, and it explains several bugs you will otherwise meet by surprise.

When one agent calls another, the calling model — the orchestrator — does not experience a colleague. It experiences a tool call. It emits a name and some arguments, the same way it would to look up an order or send an email, and at some point a result comes back as the tool's output.

Underneath, something else is happening. The sub-agent is a model too, and models take conversations. So the orchestrator's arguments have to be turned into a conversation for the sub-agent: the orchestrator's request becomes a user message, and the sub-agent's reply is its assistant message. What looked like a function call on one side is the opening turn of a conversation on the other.

That conversion is where the design decisions hide.

The sub-conversation has state. If the orchestrator reads the answer and wants to push back — "that doesn't cover the renewal case" — the honest thing is to continue the existing exchange rather than start a fresh one. That means the sub-agent's messages have to be kept in flight between calls and passed back in. Start a new conversation instead and the sub-agent answers the follow-up without the context of its own previous answer, which reads as an agent that forgets what it just said.

The orchestrator only sees what the tool returns. Whatever the sub-agent reasoned through on the way to its answer is gone unless the tool result carries it. If the orchestrator needs to know the sub-agent was unsure, uncertainty has to be part of what comes back, not a tone in prose the orchestrator has to infer.

Both sides need a shape. The sub-agent's reply travels as the content of a tool result. Give it structure — a field for the answer, a field for confidence, a field for what it couldn't do — and the orchestrator can branch on it. Leave it as free text and the orchestrator is parsing prose, which it will do inconsistently and without telling you.

How the orchestrator knows when to call

This is the question we get asked most often, and the answer has three parts.

The tool's name and description. This is the main signal, and it is a piece of writing that deserves the same care as a prompt. It should say what the sub-agent is for and, just as usefully, when not to reach for it. "Answers questions about billing policy — use for refunds, proration and invoice disputes; do not use for account access problems" routes better than "billing agent".

The parameter schema. What you ask the orchestrator to supply shapes when it thinks the tool applies. A required order_id teaches it that this tool is for cases where an order is in hand.

The system prompt. The tool description says what the tool does; the system prompt says how this assistant works. Sequencing lives here — check entitlement before offering a refund, ask for the order number before escalating — along with anything about who the assistant is that shouldn't be repeated in every tool.

When routing goes wrong, it's worth knowing which of the three to fix. A sub-agent called for cases it shouldn't touch is usually a description problem. A sub-agent called at the wrong point in a sequence is usually a system-prompt problem.

A sub-agent doesn't have to be a conversation

The sub-agent pattern doesn't imply chat. What sits behind the tool call can be:

  • One question, one answer. No follow-ups, no retained messages. Most sub-agents are this, and they're the easiest to evaluate.
  • A conversation. The orchestrator can push back, and the exchange is kept in flight for as long as the task runs.
  • An agent with its own tools. The sub-agent searches, retrieves, or calls an API on its way to answering.
  • An agent with its own sub-agents. The nesting is real, and each level costs another hop of latency, another chance to lose context, and another thing to evaluate.

Choose the simplest one that does the job. The jump from one-shot to conversational is the jump from a component you can test with a dataset to one you have to test across paths.

Evaluate every level, not just the end

A multi-agent system that scores well end to end can be wrong in ways that only show up when you look at each part. There are more things to evaluate here than people expect, and each is a different question.

Each sub-agent, on its own. Given a clean request, does it answer to the standard? This is an ordinary eval against a dataset and a rubric, and it's the one everybody does.

The orchestrator's decisions. Given a case, did it call the right sub-agent — and did it avoid calling one it shouldn't have? Both failures matter, and the second is invisible end to end when the orchestrator answers acceptably on its own.

What it does with a refusal. Sub-agents decline. When one does, does the orchestrator pass that on honestly, try a different route, or paper over it with something plausible? The last is the dangerous one and it will not show up in a happy-path test.

What it does with an error. Timeouts, malformed results, a tool that returned nothing. The orchestrator's behavior here is a design decision, and if you haven't made it deliberately, the model has made it for you.

How it reads the answer back. The sub-agent was right and the orchestrator relayed it wrong — dropped a caveat, changed a number, summarized away the condition. Evaluate the end-to-end answer against what the sub-agent actually returned, not just against the ideal.

Common mistake

Evaluating a downstream agent on upstream failures. If the orchestrator handed the sub-agent the wrong case, the sub-agent's bad answer is not the sub-agent's fault, and optimizing it will make things worse. Score each part on what it was actually given.

Watch what moves, and what it costs

Two things to keep an eye on once it runs.

The data. A run's history shows what each node received and produced, including the sub-agent's conversation. That's the difference between "the answer was wrong" and "the orchestrator called the refunds agent with the cancellation's order id."

The bill. Every hop re-sends context. Every retry multiplies. Every agent that summarizes for another pays twice for the same information. A three-step pipeline can cost several times a single well-constructed call for a marginal quality difference. Measure it against the single-agent baseline you built first, and if the multi-agent version costs four times as much and scores two points higher, make that trade deliberately rather than discovering it.

Shapes worth knowing

Three arrangements cover most of what teams actually build.

One agent, several sub-agents. The common case. An orchestrator routes to specialists and assembles the answer.

Nested. A sub-agent has sub-agents of its own. Useful when a specialist is genuinely a system — a research agent that searches, reads and summarizes — and expensive enough that it should be a deliberate choice.

Checked. Whatever the first agent produces goes to a second one that checks it against a standard before anything reaches the person. The checker is a judge, so it's the same shape Caliper scores with, and you can run it as a gate in production or as an eval in CI.

Across any of these, some parts repeat. Brand voice rules, a groundedness check, a refusal policy — written once and used by every agent that needs them, rather than copy-pasted into each and drifting apart.

Before you build the second agent

0 of 7 checked