OpenAI and Ironclad turned contracting configuration work into rubric-graded tasks for training and evaluating computer-use agents. Each task is scored on whether the finished process works across the cases it was built for. On 11 research tasks, GPT-6 Astra met 55.0% of the criteria on average, against 41.6% for GPT-5.6 Sol.

What OpenAI Published

On 6 October 2026 OpenAI published Advancing computer use with Ironclad. Ironclad is the first of a small number of software companies OpenAI is working with on this research. Ironclad employees and people who use Ironclad at OpenAI picked 11 tasks across legal, commercial and procurement work, such as setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable clause so it reflects the jurisdiction a requester selects. OpenAI estimates each task takes an experienced user 30 to 40 minutes.

Each task is graded against 8 to 50 criteria, depending on its complexity. Ironclad also provided hosted environments of its product where models could practice, and OpenAI used reinforcement learning on synthetic tasks built around representative workflows.

The Grade Is The Finished Process

OpenAI's example is a software purchasing process where Finance approves purchases above a spending threshold. The agent has to configure that rule and check that requests above and below the threshold follow the right paths. OpenAI writes that getting individual steps right is not enough, because the finished process must work across the situations it was designed to handle.

Results

  • Mean rubric score: 55.0% for GPT-6 Astra (Max reasoning) against 41.6% for GPT-5.6 Sol (High reasoning).
  • Estimated average time per attempt: 19.2 minutes for Astra against 37.0 minutes for Sol.
  • Internal model: a model used in developing Astra reached 63.7%.
  • One task side by side: Astra met about 94% of the criteria in an estimated 20 minutes, and Sol met about 85% in an estimated 32 minutes.

Ironclad's view, in OpenAI's post, is that an agent that loses track of one business rule halfway through limits what a software company can ask it to do, and that human oversight still matters as agents improve at contracting work.

What Operators Should Change

When an agent sets up approval rules, intake forms or contract templates, test the finished setup. Write one check per requirement, run a case on each side of every threshold and branch, and record which checks passed. Re-run the same checks after a model change.

What The Post Does Not Prove

  • OpenAI says the results cover the 11 research tasks, not all Ironclad workflows.
  • The times are simulated estimates based on assumed model speeds, not measured customer time savings.
  • Training tasks were built from contracts publicly available in the SEC's EDGAR database. OpenAI says it did not use customer data, its own contracts, or nonpublic Ironclad customer contracts.

Related: The verify-loop playbook covers grading agents on end state and repeated runs, including failures that look clean in the transcript. See also our note on aws-bench, an open benchmark for agents on AWS.