AI agents & workflow evaluation

Basis says GPT-6 Astra cut a 50-tab tax workbook task in half

•Make Better Editorial

Basis reports GPT-6 Astra finished a complex 50-tab tax workbook in half the time of GPT-5.6 Sol. The more useful lesson is how to evaluate long-running agents beyond token price.

Basis, which builds AI agents for accounting work, says GPT-6 Astra completed a complex 50-tab tax workbook in half the time GPT-5.6 Sol needed. The company also reports roughly a 20% improvement in its internal agent evaluation scores. Those numbers are useful, but the stronger lesson for workflow builders is how Basis thinks about long-running agent performance: task completion, judgment, corrections, reasoning effort and verification all matter alongside model price.

Evidence boundary

These results come from Basis and are published in an OpenAI customer story. They are not an independently replicated benchmark. The public material does not provide the full sample size, run-by-run accuracy results or enough methodology to generalize the 50% time reduction to other spreadsheet or accounting workloads.

What Basis reported

Reported results

50 tabs
Workbook size
Complex tax workbook used in the comparison
50% less
Completion time
GPT-6 Astra versus GPT-5.6 Sol in Basis's reported test
~20%
Internal eval improvement
Basis's own agent evaluation scores

Basis attributes part of the improvement to decisions made early in the task. According to the company, Astra takes a more direct path through the work and spends less time correcting mistakes. That can matter for agent economics because a model that has a higher per-token price can still produce a cheaper completed job if it uses fewer unnecessary steps, retries or corrections.

Why total task time is more useful than token price alone

Make Better analysis

For long-running agents, cost per million tokens is an incomplete optimization target. A production team ultimately pays for completed outcomes: model tokens, tool calls, elapsed execution, retries, failed runs and human review. A cheaper model that takes a longer path can lose its list-price advantage. Conversely, a premium model is only economically better when the reduction in corrections, execution time or human intervention is large enough to justify its higher unit cost.

Basis adjusts reasoning effort during the workflow

Basis says its agents increase reasoning effort when a step is difficult and reduce it when the work is easier. OpenAI's customer story says the model can make those changes while keeping its cache intact. The pattern is important because long tasks are rarely equally difficult from beginning to end: extraction or formatting may need little reasoning, while an ambiguous accounting decision may deserve substantially more.

The internal eval measures process, not only the final answer

Basis says its evaluation checks how the agent works as well as what it produces. Examples include whether it follows required templates, consults primary sources for tax questions and checks its own work. The company reports about a 20% improvement in these internal evaluation scores with Astra and attributes part of that gain to better understanding of user intent.

Make Better analysis

This is a useful design principle even if another team never uses Astra. Agent evaluations should include trajectory requirements that matter to the business. A correct-looking final document is not enough if the agent skipped a required source, ignored an approval step or created hidden rework for a human reviewer.

What the public case study does not prove

  • It does not establish that Astra will cut every long-running agent task by 50%.
  • The public material does not disclose a full sample size or repeated-run distribution for the workbook comparison.
  • It does not publish enough accuracy and correction data to calculate a verified cost per correct completed workbook.
  • The ~20% result refers to Basis's internal evaluation, so it should not be compared directly with a public benchmark score.
  • Results from a structured accounting workflow may not transfer to research, coding, sales or marketing agents.

A better scorecard for long-running agents

Measure the completed job

MetricWhat it tells you
Successful completion rateHow often the agent reaches the required outcome without a restart
Total task costTokens, tool calls and other execution costs for a completed job
Elapsed task timeWhether the workflow is operationally fast enough
Human correction timeHow much hidden labor remains after the agent finishes
Retry / recovery rateHow often the agent takes a wrong path or needs another attempt
Process complianceWhether required sources, templates, approvals and checks were actually used

How to test a model upgrade in your own workflow

  1. Choose a real recurring task with a clear definition of successful completion.
  2. Run the current and candidate models on the same representative task set.
  3. Record total execution time, tokens, tool calls, retries and human corrections.
  4. Score process compliance separately from final-output quality.
  5. Calculate cost per successful completed task rather than comparing token price alone.
  6. Route only the workloads that show a meaningful measured benefit to the more expensive model.
Bottom line

Basis's reported 50-tab workbook result is interesting evidence, not a universal benchmark. The reusable lesson is stronger: evaluate agents as end-to-end workers. Measure whether they finish correctly, how directly they get there, what they cost after retries and review, and whether they follow the process your business actually requires.

Sources & useful resources