Subagents Explained: How AI Agents Delegate Work to Other Agents
Learn what subagents are, how AI agents delegate tasks, and when multi-agent workflows outperform a single agent, with architecture patterns, risks, and design guidance.
Subagents Explained: How AI Agents Delegate Work to Other Agents
Anthropic reported that its multi-agent research system outperformed a single-agent setup by 90.2% on an internal research evaluation, while using roughly 15 times the tokens of ordinary chat interactions. That combination captures the promise and the price of subagents: better coverage and parallel work, but only when delegation is worth the extra coordination and compute (Anthropic, 2025).
Subagents are separately operating AI agents given bounded assignments by a coordinating agent. Each receives instructions, context, tools, and an output contract, then returns its work to the coordinator for validation and synthesis.
The important idea is delegation, not headcount. Adding agents does not automatically make a system smarter. Good delegation creates useful specialization, context isolation, and parallelism. Bad delegation creates duplicated research, conflicting answers, runaway cost, and a much harder system to debug.
A subagent is not a smaller intelligence living inside another agent. It is a bounded execution context that owns one delegated piece of a larger outcome.
Key takeaways
- Subagents handle scoped parts of a larger task under a parent or orchestrator agent.
- They can use the same underlying model or different models, tools, permissions, and context.
- Parallel subagents help only when workstreams are meaningfully independent.
- The coordinating agent remains accountable for validation, synthesis, and stopping conditions.
- A single agent, script, API call, or workflow is often the better design.
- Production subagent systems need traces, budgets, least privilege, and human approval for consequential actions.
What is a subagent?
A subagent is an AI agent invoked or created to complete a specific assignment within a larger task. A parent, lead, or orchestrator agent owns the overall goal. It decides what to delegate, supplies the relevant context, and evaluates the result.
The labels vary across products:
| Role | Typical responsibility |
|---|---|
| Parent or lead agent | Owns the final outcome and user-facing response |
| Orchestrator | Plans, assigns, monitors, and combines work |
| Subagent | Completes one bounded delegated assignment |
| Specialist agent | Handles a domain or toolset with tailored instructions |
| Worker agent | Executes a general task under an orchestrator |
There is no universal technical specification for a subagent. Some implementations support parallel execution, nested delegation, or persistent memory. Others do not. The stable concept is that responsibility is delegated and results return to a coordinator.
How a subagent differs from a model call
A normal model call maps an input to an output. A subagent may work through multiple steps, call tools, maintain temporary state, adapt its approach, and return a structured artifact. It has a goal to pursue within boundaries, not merely a prompt to complete once.
That distinction also separates subagents from background jobs. Asynchronous execution alone is not agentic. A scheduled database export is still a job. A worker that can interpret a goal, choose sources, use tools, recover from errors, and stop when evidence is sufficient behaves more like a subagent.
How subagent delegation works
The basic lifecycle is simple:
User request → lead agent → task plan → subagents A, B, and C → returned results → validation and synthesis → final output
The hard part is designing each boundary.
1. The lead agent interprets the outcome
The coordinator identifies the requested outcome, dependencies, evidence standards, risks, and definition of done. It should ask whether the task truly benefits from delegation before creating workers.
2. It decomposes the work into bounded assignments
A useful assignment specifies:
“Research competitors” is vague. “Identify five direct competitors, verify public pricing on primary pages, record the access date, and return a comparison table with citations” is bounded and testable.
For a reusable version, this compact delegation brief gives an agent much less room to wander:
{ "objective": "Verify public pricing for five direct competitors", "scope": { "include": ["official pricing pages", "plan names", "monthly price", "annual billing terms"], "exclude": ["review-site estimates", "unverified forum posts"] }, "tools": ["web search", "web fetch"], "output": { "format": "table", "fields": ["company", "plan", "price", "billing period", "source_url", "checked_at"] }, "budget": { "max_sources": 15, "max_minutes": 12 }, "done_when": "Every claim has a primary source or is marked not publicly available", "escalate_when": "Sources conflict or require authenticated access" }3. Subagents execute sequentially or in parallel
Sequential delegation fits dependent work, such as research followed by drafting and then review. Parallel delegation fits independent branches, such as checking pricing, product capabilities, and customer sentiment at the same time.
Parallel execution can reduce elapsed time, but not total compute. Anthropic reported research-time reductions of up to 90% for complex queries in its implementation, while also reporting substantially higher token use. Those are internal results, not a universal benchmark.
4. Results return to the orchestrator
The coordinator checks whether the assignment was completed, verifies evidence, resolves disagreements, removes duplication, and requests targeted follow-up work. It then synthesizes the final answer or action.
5. The orchestrator retains accountability
Delegation does not transfer final responsibility. The lead agent should still own permissions, stopping conditions, source quality, and the final result. For high-impact actions, it should pause for human approval rather than letting a worker act merely because it can.
Subagents vs tools, workflows, handoffs, and multi-agent systems
These concepts overlap, but they are not interchangeable.
| Concept | Autonomy | Own context | Multiple tool calls | Usually returns control |
|---|---|---|---|---|
| Tool | Low | No | No | Immediately |
| Workflow step | Usually low | Sometimes | Sometimes | By design |
| Subagent | Medium to high | Usually | Often | Usually |
| Handoff agent | Medium to high | Often | Often | Not necessarily |
| Peer agent | Medium to high | Usually | Often | No fixed parent |
Implementations differ. This table describes common patterns, not a formal standard.
Subagents vs tools
A tool exposes a capability such as search, email, code execution, or a database query. A subagent decides how to pursue a delegated objective, often through several tool calls. An agent can use tools. A tool is not automatically an agent.
Subagents vs workflows
A conventional workflow follows predefined steps. A subagent can adapt its approach while working. Good production systems often combine both: deterministic software handles known transformations, while agents handle ambiguous judgment.
This is why a flowchart is not the same thing as delegation. Zapier, Make, and n8n are useful when inputs, branches, and actions can be mapped in advance. Subagents fit work where the route must change as evidence appears. Replacing a stable API call with three agents is not sophistication. It is waste.
Subagents vs multi-agent systems
Subagents are one pattern inside the broader category of multi-agent systems. Not every multi-agent architecture has a parent-child hierarchy. Agents can collaborate through peer-to-peer exchange, debate, voting, routing, or handoffs.
Subagents vs handoffs
With delegation, the lead agent usually stays in control and expects a result back. With a handoff, another agent may take over the conversation or responsibility. The OpenAI Agents SDK documentation distinguishes a manager calling agents as tools from handoffs that transfer control.
Why AI systems use subagents
Parallelism
Independent research or analysis branches can run at the same time. This matters when breadth is valuable and wall-clock time is constrained.
Context isolation
Each worker receives only the material relevant to its assignment. That reduces irrelevant context and cross-task interference. The coordinator gets compressed findings instead of every intermediate observation.
Specialization
Subagents can have different instructions, models, tools, permissions, knowledge sources, and output schemas. This is configuration, not proof that the underlying model has become a distinct intelligence.
Broader exploration
Several workers can investigate different hypotheses, markets, sources, or implementation paths. A single sequential agent is more likely to stay on its first plausible trajectory.
Fault containment
A failed bounded assignment can sometimes be retried without discarding the whole run. That benefit requires checkpoints and orchestration. It does not happen automatically.
How we built our multi-agent research systemOn the the engineering challenges and lessons learned from building Claude's Research systemA practical example: building a competitive market report
Suppose a lead agent must produce a sourced market report. Instead of asking four generic workers to “research the market,” it creates non-overlapping assignments:
| Subagent | Assignment | Required output |
|---|---|---|
| Market sizing | Find credible category estimates and methodology | Claims table with source quality notes |
| Product research | Compare official product capabilities | Capability matrix with primary URLs |
| Pricing verification | Check official prices and billing terms | Normalized pricing table with access dates |
| Review analysis | Analyze a defined sample of customer reviews | Themes, counts, quotations, limitations |
The lead agent still must reconcile category definitions, reject weak estimates, check citations, remove duplicate claims, and write the report. If two workers disagree, it needs an explicit source hierarchy rather than a majority vote.
The same pattern applies elsewhere:
- In software development, workers can explore a codebase, implement a change, run tests, and review the patch.
- In marketing, they can inspect search intent, analyze competing pages, draft, and fact-check.
- In operations, they can collect records, validate exceptions, and prepare an approval queue.
The architecture helps only when the assignments are independently useful and their outputs can be evaluated.
Five common subagent architecture patterns
| Pattern | Best use | Advantage | Main risk |
|---|---|---|---|
| Orchestrator-worker | Several bounded workstreams | Clear ownership | Coordinator becomes a bottleneck |
| Router-specialist | One request needs one expert | Efficient specialization | Wrong routing decision |
| Parallel fan-out | Independent research branches | Faster broad coverage | Duplicate work and contradictions |
| Sequential pipeline | Output needs staged transformation | Clear processing order | Compounding errors |
| Reviewer agent | Draft needs a second pass | Consistent checks | False confidence from AI reviewing AI |
Microsoft documents the orchestrator and subagent pattern as a coordinator delegating to specialized agents and aggregating their outputs. Google Cloud's multi-agent architecture guidance similarly emphasizes communication, security, deployment, and observability in production systems. These patterns are architectural choices, not proof that every step should be agentic.

AI Agent Orchestration: How to Coordinate Multi-Agent Systems
AI agent orchestration coordinates multiple reasoning agents toward one goal. Learn the four patterns, the architecture, and how to run a multi-agent system with no code.
When subagents are the right choice
Use subagents when a task:
- Splits into clear, independent workstreams.
- Requires broad research across many sources.
- Exceeds one context window's practical capacity.
- Benefits from different tools, permissions, or specialties.
- Has enough value to justify additional model usage.
- Produces outputs that can be objectively checked and merged.
Good examples include due diligence, large-codebase exploration, multi-market localization, comparative analysis, multi-source monitoring, and draft-plus-review processes.
When not to use subagents
A single agent or deterministic system is usually better when:
- The task is small or linear.
- Every step depends tightly on the previous step.
- All workers need the same complete context.
- Latency and token use matter more than breadth.
- The output cannot be evaluated reliably.
- A database query, API call, or script already solves the problem.
- Coordination costs more than the task is worth.
Start with the simplest architecture that can complete the task reliably. Add subagents only when measured limitations justify them.
A chatbot also does not become an operational system merely because it can describe a plan. Chat is an interface. Reliable delegation requires scoped execution, permissions, traces, evaluation, and clear control over external actions.
| Choose | When it fits |
|---|---|
| Zapier, Make, or n8n | Steps, branches, and actions are known in advance |
| A chatbot | The user primarily needs an interactive answer |
| One agent | The task is ambiguous but mostly linear |
| Subagents | Independent workstreams need adaptive reasoning, separate context, or parallel execution |
| Rerun | Agents must run continuously with connected tools, monitoring, logs, and approvals |
Risks and failure modes
Duplicate work and conflicting answers
Overlapping assignments waste tokens and can leave important gaps. Define ownership before execution, then give the coordinator source-quality and conflict-resolution rules.
Context loss
A worker may miss background obvious to the parent. Passing everything defeats isolation. Send the smallest complete context and make assumptions explicit.
Compounding errors
Sequential systems can amplify an early mistake. Preserve evidence, validate stage outputs, and avoid treating earlier prose as ground truth.
Uncontrolled cost and latency
More workers mean more prompts, model calls, and tool use. Set maximum agent counts, token budgets, deadlines, retry limits, and cancellation rules.
Excessive permissions
Each worker should have the minimum tools and data needed. A research worker does not need permission to publish, transfer money, or modify customer records.
Harder evaluation and debugging
Nondeterministic paths require traces showing which agent ran, what it received, which tools it called, what it returned, and why the coordinator accepted or rejected the result.
How to design reliable subagent workflows
- Give every worker one responsibility. Avoid vague prompts and overlapping ownership.
- Define the output before execution. Prefer schemas, tables, citations, files, or checklists over unconstrained prose.
- Restrict tools and permissions. Apply least privilege and approval gates.
- Set budgets and stopping conditions. Limit depth, retries, tool calls, tokens, and elapsed time.
- Preserve provenance. Keep primary sources attached to claims and label inferences.
- Evaluate outcomes, not one path. Measure success, accuracy, completeness, cost, latency, redundant work, and human intervention.
- Keep humans in the loop. Require approval for publishing, transactions, account changes, sensitive data access, and other consequential actions.
Rerun fits at the operations layer, not as a multi-agent framework. It gives each agent an always-on private cloud machine, connected tools, schedules, action logs, monitoring, and human approval gates. It does not replace orchestration libraries or define how a developer should implement recursive subagent logic. The value is making ongoing agent work visible and controllable once it touches real operations.
The bottom line
Subagents are a delegation mechanism, not an automatic intelligence multiplier. They work best when a valuable task separates into bounded workstreams that can run independently and be objectively checked.
The winning design is rarely the one with the most agents. It is the one with clear responsibility, narrow permissions, verifiable outputs, visible traces, sensible budgets, and a human decision point where consequences matter.
Frequently asked questions
What are subagents?
Subagents are separately operating AI agents given bounded assignments by a parent or orchestrator agent. They receive instructions, context, tools, and an output contract, then return their work for validation and synthesis.
Are subagents separate AI models?
Not necessarily. Several subagents can use the same model with different instructions and context, or use different models selected for capability, speed, or cost.
Do subagents share memory?
It depends on the implementation. Some share selected state through an orchestrator, while others work in isolated contexts. Persistent shared memory should never be assumed.
Can a subagent create another subagent?
Some systems allow nested delegation and others impose depth limits. Nesting can expand coverage, but it also raises coordination, cost, permission, and debugging risks.
Are subagents faster than one agent?
They can reduce elapsed time when independent tasks run in parallel. They may still consume more total compute and can be slower when workstreams depend tightly on one another.
How many subagents should a system use?
There is no universal number. Use one worker per genuinely independent, valuable workstream, then enforce agent-count, cost, time, and retry limits.
Are subagents the same as a multi-agent system?
Subagents are one multi-agent pattern, usually implying delegated responsibility under a parent. Multi-agent systems can also use peers, routers, debates, voting, or handoffs without a strict hierarchy.
When should you not use subagents?
Avoid them for small or linear tasks, work that requires identical full context, outputs that cannot be evaluated, or problems already solved reliably by a script, API call, database query, or deterministic workflow.
Written by
Clément Janssens

