Anthropic is partnering with Accenture on independent embedded evaluation of frontier AI — evaluating and red-teaming models, running alignment assessments, and testing safeguards. This is a capacity-and-access arrangement, not a new public scorecard and not a restatement of Anthropic's pace-of-development measurements.

What Anthropic Published

On September 18, 2026, Anthropic described a partnership with Accenture on embedded evaluation. The work will be led by Faculty, Accenture's specialist AI business. Anthropic frames the partnership as a step toward the commitment in Dario Amodei's essay "We Must Pace the Frontier": embed evaluators inside Anthropic with access comparable to an employee.

Accenture deploys AI with businesses and governments across industries. Anthropic says that enterprise-use perspective informs Accenture's safety approach and will inform how Faculty evaluates Anthropic's models. This post names a partner and a work scope. It does not publish evaluation results, incident reports, or Institute metrics.

What Embedded Evaluation Is Meant To Do

Anthropic says embedded evaluation is new, and many operating details are still being worked out. Unlike today's external evaluators, embedded evaluators would work inside AI companies with access comparable to an employee. That access is meant to let them watch models take shape in training, follow the decisions that govern how those models are built and deployed, and speak directly to employees.

From that vantage point, Anthropic says embedded evaluators can:

  • Assess how a company operates.
  • Verify that it is keeping its safety commitments.
  • Identify blind spots.
  • Report incidents and give the public a more informed account of benefits and risks.

The named work includes evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards. Anthropic does not claim those activities already have settled methods or shared public outputs.

Funding, Non-Exclusivity, And The Evaluator Ecosystem

Anthropic and Accenture each expect to invest at least $1 billion building capacity in this area over the next five years. Anthropic will fund Accenture's work directly for now. It is also in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation under other funding arrangements — in those cases, using the evaluators' own funding.

The partnership is non-exclusive. Anthropic says it expects frontier labs to work with several organizations at once, that it will work with other evaluators to be announced in the coming weeks, and that Accenture will work with other AI developers in similar capacities. Long-term, Anthropic argues frontier AI needs an ecosystem of evaluators operating with shared standards, and that funding should come from pooled or government sources as it called for in its Advanced AI Framework in June. It says neither of those funding sources exists today.

Accountability Stays With Anthropic

Anthropic states that independent embedded evaluators do not reduce its accountability. The safety of its models remains Anthropic's responsibility. Evaluators are described as making that responsibility more verifiable, not transferring it.

There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation. Anthropic says it will continue to train and release frontier models, wants independent evaluators working alongside that work, and expects the approach to evolve as the field matures.

What The Announcement Does Not Prove

  • A named partner and a capacity commitment are not published evaluation results or a public incident log.
  • Employee-comparable access is the intended design. The post says operating details are still being worked out.
  • The $1 billion / five-year figures are each side's expected investment in capacity. They are not an Institute score and not a completed spend.
  • This is distinct from Anthropic's pace-of-development measurements (AI-led R&D share, agent oversight, and safety-compute share). Those are lab-internal process metrics, not this evaluator arrangement.
  • Embedded evaluators do not move safety off Anthropic's books. The post keeps accountability with the lab.

Related: See our notes on Anthropic's pace-of-development measurements and Anthropic's Enterprise Frontier Safeguards.