Anthropic has published three lab-internal measurements for tracking how fast frontier AI is being built: how much of its R&D Claude performs, how those agents are overseen, and how compute is split between safety and other research. This is a snapshot of production process, not a disclosure taxonomy for individual misalignment cases.
What Anthropic Published
On September 17, 2026, Anthropic described measurements for understanding the pace of AI development inside frontier labs. The post argues that capability evaluations alone do not show how models are produced, and that the public needs visibility into those inputs as governments consider pacing the frontier.
The three measurements are (1) the share of Anthropic AI R&D performed by Claude, (2) oversight of agents acting on Anthropic's systems, and (3) how compute is allocated, including the share that goes to safety. Anthropic presents a current snapshot and says it would expect the numbers to move if labs coordinated on pacing. It also says it plans to embed independent third-party evaluators with access comparable to internal risk teams. This is a different object from OpenAI's standing process for disclosing individual misalignment cases.
AI-Led R&D: Leads, Collaborates, Not Autonomous
Anthropic built a prototype Anthropic R&D Automation Index by cataloguing kinds of AI R&D work, rating how automated each task is, and aggregating those ratings. The ratings use Epoch AI's Automation Level scale. AL3 ("collaborates") means the model can do large chunks of work under close human direction. AL4 ("leads") means it can complete most of a task end-to-end from a high-level prompt while a human supervises. AL5 means the model operates fully autonomously, with no human in the loop.
As of August 2026, Anthropic reports:
- Claude is not operating fully autonomously for any measured subset of AI R&D work (no AL5).
- Claude "leads" 26% of Anthropic's AI R&D work.
- The share of work at or above "AI collaborates" is above 90%.
The company says any frontier developer could publish the same measures on a public methodology, and that cross-lab comparison still lacks a shared method and independent judges. It does not claim these figures are a public Institute score or a forecast of when AL5 arrives.
Oversight Of Internal Agents
Anthropic says employees increasingly delegate long-running work to agents that also delegate to one another. As work moves from collaboration toward leadership, those agents can make more consequential research choices. The post distinguishes harms that must be stopped in real time from slower patterns that can be reviewed after the fact.
As of August 2026, approximately 30,000 agents were doing research and engineering work at Anthropic at any one time on its most-used internal platform. The measurements cover that platform only. Anthropic analyzed over a billion decisions from those research and engineering agents in August 2026 and reports that 0.002% of them were blocked by the online monitor.
The post also describes online monitors (real-time block or steer) and offline monitors (after-the-fact review). It does not publish a complete public effectiveness claim, and it notes that METR has previously red-teamed the offline monitoring platform. Coverage and review-latency figures in the post are platform-specific and are not restated here as operator guarantees.
Safety Compute Share
Anthropic examined a snapshot of how it used all of its compute from July 13 to July 20. Over that week, about 6% of compute that went to AI R&D was allocated toward safety, and about 12% of compute that went to AI-driven AI R&D was allocated toward safety.
Those estimates are deliberately conservative. Work that advanced capabilities as much as safety was not counted as safety, and the metrics exclude safeguards classifiers, which Anthropic says are a separate, comparable amount of compute. The company also notes that safety research is often less compute-intensive than frontier training, so the share is an imperfect proxy for focus. The value it claims is comparability across developers and over time, not an absolute grade.
What The Snapshot Does Not Prove
- 26% "leads" is not recursive self-improvement. Anthropic reports no AL5 on the measured work.
- A one-week compute share is a snapshot of a fungible pool, not a fixed budget or a safety-work count.
- ~30,000 concurrent agents and a 0.002% online block rate describe one internal platform, not the whole economy of agents.
- Publishing lab-internal pace metrics is not the same as disclosing individual misalignment incidents.
- These figures are Anthropic's own measurements. They are not Institute scores, and they are not a price or trading signal.
Related: See our notes on OpenAI's model-misalignment reporting framework and Anthropic's Enterprise Frontier Safeguards.