Computer-use agents & evaluation

OpenAI and Ironclad show how to evaluate computer-use agents on real workflows

•Make Better Editorial

OpenAI and Ironclad turned 11 contracting workflows into detailed agent evaluations. The useful lesson is a practical framework for testing computer-use agents before trusting them with business-critical work.

OpenAI and Ironclad published a research collaboration on October 6 that turns complex contracting work into training and evaluation tasks for computer-use agents. The important part for builders is not only that GPT-6 Astra outperformed GPT-5.6 Sol. The collaboration shows a more useful way to test agents: define realistic end-to-end workflows, score them against many explicit requirements, and give models a safe environment where they can practice and be evaluated.

What the evaluation tested

Ironclad employees and people who use Ironclad at OpenAI identified 11 tasks spanning legal, commercial and procurement work. Examples included configuring nondisclosure agreements, creating procurement approval processes and updating reusable legal clauses based on a requester’s jurisdiction. OpenAI says an experienced user would take roughly 30 to 40 minutes per task on average.

Research evaluation

11
Workflows
Legal, commercial and procurement tasks
8–50
Criteria per task
Depending on task complexity
55.0%
GPT-6 Astra
Mean rubric score
41.6%
GPT-5.6 Sol
Mean rubric score

Why rubric design matters more than a single success flag

Make Better analysis

A business workflow can look complete while still violating one important rule. An agent might create the right intake form but miss a finance threshold, route an exception incorrectly, or fail to preserve a reusable legal requirement. Scoring 8 to 50 criteria per task exposes partial failures that a simple pass/fail benchmark can hide. For teams deploying agents, this suggests evaluating the finished business state and its constraints—not just whether the agent clicked the right buttons.

A practical evaluation framework for business agents

  1. Choose a small set of high-value workflows that represent real production work, including exceptions rather than only happy paths.
  2. Ask domain experts to define the requirements that must remain true at the end of each workflow.
  3. Turn those requirements into a granular rubric so partial success and silent failures are visible.
  4. Run every model or agent configuration in the same controlled environment with the same task inputs.
  5. Measure task quality, retries, human corrections, latency and cost together instead of optimizing a single benchmark score.
  6. Keep human approval around consequential actions until the agent consistently passes the criteria that matter to the business.
  7. Retest after model, prompt, tool or workflow changes because agent reliability belongs to the whole system, not only the underlying model.

What the Astra numbers do—and do not—show

Across the 11 research tasks, OpenAI reports a 55.0% average rubric score for GPT-6 Astra at Max reasoning versus 41.6% for GPT-5.6 Sol at High reasoning. Estimated average time per attempt fell from 37.0 minutes to 19.2 minutes. OpenAI explicitly cautions that these times are simulated estimates based on assumed model processing and generation speeds. They are not measured customer productivity gains, and the 11 tasks do not represent every Ironclad workflow.

Evidence boundary

Treat the score improvement as evidence on this specific research evaluation, not as a general 32% productivity gain. Production ROI also depends on task mix, human review, retries, error severity, integration overhead and the cost of failures.

Bottom line

The strongest lesson from the OpenAI–Ironclad work is methodological: evaluate agents on complete workflows with detailed, domain-specific criteria and controlled environments. Model capability is improving, but reliable deployment requires measuring whether the final business process still obeys every rule that matters.

Sources & useful resources