← Back to Playbooks
Engineering

Multi-Agent Research Cost Control

Scale agent count to query complexity — multi-agent is a 15× token gate, not a default

Operator recipe for when a lead agent fans out parallel research subagents. Scale agent count and tool calls to query complexity, teach the orchestrator to delegate with objectives and boundaries (avoid duplicate searches), prefer parallel tool calls for speed, and treat token burn as an economic gate. Anthropic’s engineering post (published 13 Jun 2025 on the page) reports: a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on an internal research eval; three factors explained 95% of BrowseComp performance variance, with token usage by itself explaining 80%; agents typically use about 4× more tokens than chat interactions, and multi-agent systems about 15× more than chats. Reserve multi-agent for high-value, highly parallelizable work. Same primary also documents production reliability: resume-from-checkpoint on errors, rainbow deployments for long-running stateful agents, and subagent artifacts / filesystem handoffs to reduce telephone loss. Parallelization cut research time by up to 90% for complex queries; a tool-description rewrite agent produced a 40% decrease in task completion time. Those figures are Anthropic’s, not Institute metrics. Not a Brief thesis swap. Distinct from /playbooks/verify-loop-ci-agent-code/ (CI test selection when agents write code) and /playbooks/harness-permission-outside-agent/ (host grants connectors/egress). Primary: https://www.anthropic.com/engineering/multi-agent-research-system.

0
Registrations
$1,000
Prize Pool
Feb 23, 2026
Starts
0%
Complete

Core Workflows (6)

Each workflow represents a critical business function. Click any workflow to see detailed automation architecture.

01

Apply the economic gate before you fan out

Led by: Research Operator

Anthropic: agents typically use about 4× more tokens than chat; multi-agent systems about 15× more than chats. For economic viability, multi-agent requires tasks where the value of the task is high enough to pay for the increased performance. Multi-agent systems excel at valuable tasks that involve heavy parallelization, information that exceeds single context windows, and interfacing with numerous complex tools. They are a poor fit today when all agents must share the same context or when there are many dependencies between agents — Anthropic names most coding tasks as having fewer truly parallelizable units than research. Token usage explained 80% of BrowseComp variance (tokens + tool calls + model choice = 95%). Cite those figures as Anthropic’s, not Institute stats. Do not default multi-agent because a lead agent can spawn subagents.

Sub-Agents
Lead Researcher
Skills Required
Token economicsTask-value gate
Human TouchpointOnly authorize multi-agent when the task value covers ~15× chat tokens and the work is actually parallelizable
02

Scale agent count and tool calls to query complexity

Led by: Lead Researcher

Anthropic embedded scaling rules in prompts because agents struggle to judge effort. Simple fact-finding: 1 agent with 3–10 tool calls. Direct comparisons: 2–4 subagents with 10–15 calls each. Complex research: more than 10 subagents with clearly divided responsibilities. Early failure mode: spawning 50 subagents for simple queries, or overinvesting in fact-finding. The 90.2% internal-eval lift (Opus 4 lead + Sonnet 4 subagents vs single Opus 4) is a breadth-first research result — not a reason to use that shape for a one-fact lookup.

Sub-Agents
Subagent
Skills Required
Effort scalingQuery complexity
Human TouchpointRefuse a 10-subagent plan for a fact-finding query; match agent count to the complexity band
03

Teach the orchestrator to delegate with boundaries

Led by: Lead Researcher

The lead agent decomposes queries into subtasks. Each subagent needs an objective, an output format, guidance on tools and sources, and clear task boundaries. Without detailed task descriptions, agents duplicate work, leave gaps, or run the same searches. Anthropic’s example: a vague “research the semiconductor shortage” sent one subagent into the 2021 automotive chip crisis while two others duplicated 2025 supply-chain searches. Start wide, then narrow (short, broad queries first). Distinct from /playbooks/verify-loop-ci-agent-code/ — this is research orchestration, not CI test selection.

Sub-Agents
Subagent
Skills Required
Orchestrator delegationTask boundaries
Human TouchpointReview the first fan-out: each subagent should have a distinct objective and source set, not a shared vague brief
04

Prefer parallel tool calls for speed

Led by: Lead Researcher

Anthropic introduced two kinds of parallelization: (1) the lead agent spins up 3–5 subagents in parallel rather than serially; (2) the subagents use 3+ tools in parallel. These changes cut research time by up to 90% for complex queries. Separately, a tool-testing agent that rewrote a flawed MCP tool description produced a 40% decrease in task completion time for later agents using the new description. Sequential search was the early slow path. Do not treat the 90% / 40% figures as Institute throughput.

Sub-Agents
Subagent
Skills Required
Parallel tool callsTool-description hygiene
Human TouchpointPrefer 3–5 parallel subagents and 3+ parallel tools on complex queries; fix tool descriptions before blaming the model
05

Resume from checkpoint; rainbow-deploy stateful agents

Led by: Reliability Operator

Agents are stateful and errors compound. Anthropic: when errors occur, do not restart from the beginning — restarts are expensive. Build systems that resume from where the agent was, combined with retry logic and regular checkpoints. Deployment: you cannot update every long-running agent to a new version at the same time. Use rainbow deployments — gradually shift traffic from old to new versions while keeping both running — so a well-meaning code change does not break agents mid-process. Distinct from /playbooks/harness-permission-outside-agent/ (egress/approval), which is not a deploy strategy.

Sub-Agents
Lead Researcher
Skills Required
CheckpointsRainbow deployments
Human TouchpointDo not kill in-flight research sessions on a prompt/tool deploy; drain via rainbow cutover
06

Hand off artifacts, not a game of telephone

Led by: Lead Researcher

Anthropic appendix: subagent output to a filesystem minimizes the “game of telephone.” Subagents store work in external systems and pass lightweight references back to the coordinator, instead of copying large outputs through conversation history. The lead agent also saves its plan to Memory so the plan survives context-window truncation (the post notes a 200,000-token window). When context limits approach, agents can spawn fresh subagents with clean contexts while retrieving the stored plan. Use this for structured outputs (code, reports, visualizations) where the subagent’s specialized prompt beats filtering through a general coordinator.

Sub-Agents
SubagentCitation Agent
Skills Required
Artifact storeMemory / plan persistence
Human TouchpointRequire filesystem or artifact references from subagents; do not pipe full reports through the lead transcript