Multi-Agent Research Cost Control
Scale agent count to query complexity — multi-agent is a 15× token gate, not a default
Operator recipe for when a lead agent fans out parallel research subagents. Scale agent count and tool calls to query complexity, teach the orchestrator to delegate with objectives and boundaries (avoid duplicate searches), prefer parallel tool calls for speed, and treat token burn as an economic gate. Anthropic’s engineering post (published 13 Jun 2025 on the page) reports: a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on an internal research eval; three factors explained 95% of BrowseComp performance variance, with token usage by itself explaining 80%; agents typically use about 4× more tokens than chat interactions, and multi-agent systems about 15× more than chats. Reserve multi-agent for high-value, highly parallelizable work. Same primary also documents production reliability: resume-from-checkpoint on errors, rainbow deployments for long-running stateful agents, and subagent artifacts / filesystem handoffs to reduce telephone loss. Parallelization cut research time by up to 90% for complex queries; a tool-description rewrite agent produced a 40% decrease in task completion time. Those figures are Anthropic’s, not Institute metrics. Not a Brief thesis swap. Distinct from /playbooks/verify-loop-ci-agent-code/ (CI test selection when agents write code) and /playbooks/harness-permission-outside-agent/ (host grants connectors/egress). Primary: https://www.anthropic.com/engineering/multi-agent-research-system.
Core Workflows (6)
Each workflow represents a critical business function. Click any workflow to see detailed automation architecture.
Apply the economic gate before you fan out
Led by: Research OperatorAnthropic: agents typically use about 4× more tokens than chat; multi-agent systems about 15× more than chats. For economic viability, multi-agent requires tasks where the value of the task is high enough to pay for the increased performance. Multi-agent systems excel at valuable tasks that involve heavy parallelization, information that exceeds single context windows, and interfacing with numerous complex tools. They are a poor fit today when all agents must share the same context or when there are many dependencies between agents — Anthropic names most coding tasks as having fewer truly parallelizable units than research. Token usage explained 80% of BrowseComp variance (tokens + tool calls + model choice = 95%). Cite those figures as Anthropic’s, not Institute stats. Do not default multi-agent because a lead agent can spawn subagents.
Scale agent count and tool calls to query complexity
Led by: Lead ResearcherAnthropic embedded scaling rules in prompts because agents struggle to judge effort. Simple fact-finding: 1 agent with 3–10 tool calls. Direct comparisons: 2–4 subagents with 10–15 calls each. Complex research: more than 10 subagents with clearly divided responsibilities. Early failure mode: spawning 50 subagents for simple queries, or overinvesting in fact-finding. The 90.2% internal-eval lift (Opus 4 lead + Sonnet 4 subagents vs single Opus 4) is a breadth-first research result — not a reason to use that shape for a one-fact lookup.
Teach the orchestrator to delegate with boundaries
Led by: Lead ResearcherThe lead agent decomposes queries into subtasks. Each subagent needs an objective, an output format, guidance on tools and sources, and clear task boundaries. Without detailed task descriptions, agents duplicate work, leave gaps, or run the same searches. Anthropic’s example: a vague “research the semiconductor shortage” sent one subagent into the 2021 automotive chip crisis while two others duplicated 2025 supply-chain searches. Start wide, then narrow (short, broad queries first). Distinct from /playbooks/verify-loop-ci-agent-code/ — this is research orchestration, not CI test selection.
Prefer parallel tool calls for speed
Led by: Lead ResearcherAnthropic introduced two kinds of parallelization: (1) the lead agent spins up 3–5 subagents in parallel rather than serially; (2) the subagents use 3+ tools in parallel. These changes cut research time by up to 90% for complex queries. Separately, a tool-testing agent that rewrote a flawed MCP tool description produced a 40% decrease in task completion time for later agents using the new description. Sequential search was the early slow path. Do not treat the 90% / 40% figures as Institute throughput.
Resume from checkpoint; rainbow-deploy stateful agents
Led by: Reliability OperatorAgents are stateful and errors compound. Anthropic: when errors occur, do not restart from the beginning — restarts are expensive. Build systems that resume from where the agent was, combined with retry logic and regular checkpoints. Deployment: you cannot update every long-running agent to a new version at the same time. Use rainbow deployments — gradually shift traffic from old to new versions while keeping both running — so a well-meaning code change does not break agents mid-process. Distinct from /playbooks/harness-permission-outside-agent/ (egress/approval), which is not a deploy strategy.
Hand off artifacts, not a game of telephone
Led by: Lead ResearcherAnthropic appendix: subagent output to a filesystem minimizes the “game of telephone.” Subagents store work in external systems and pass lightweight references back to the coordinator, instead of copying large outputs through conversation history. The lead agent also saves its plan to Memory so the plan survives context-window truncation (the post notes a 200,000-token window). When context limits approach, agents can spawn fresh subagents with clean contexts while retrieving the stored plan. Use this for structured outputs (code, reports, visualizations) where the subagent’s specialized prompt beats filtering through a general coordinator.