Basis says GPT-6 Astra cut a 50-tab tax workbook task in half
Basis reports GPT-6 Astra finished a complex 50-tab tax workbook in half the time of GPT-5.6 Sol. The more useful lesson is how to evaluate long-running agents beyond token price.
Basis, which builds AI agents for accounting work, says GPT-6 Astra completed a complex 50-tab tax workbook in half the time GPT-5.6 Sol needed. The company also reports roughly a 20% improvement in its internal agent evaluation scores. Those numbers are useful, but the stronger lesson for workflow builders is how Basis thinks about long-running agent performance: task completion, judgment, corrections, reasoning effort and verification all matter alongside model price.
These results come from Basis and are published in an OpenAI customer story. They are not an independently replicated benchmark. The public material does not provide the full sample size, run-by-run accuracy results or enough methodology to generalize the 50% time reduction to other spreadsheet or accounting workloads.
What Basis reported
Reported results
Basis attributes part of the improvement to decisions made early in the task. According to the company, Astra takes a more direct path through the work and spends less time correcting mistakes. That can matter for agent economics because a model that has a higher per-token price can still produce a cheaper completed job if it uses fewer unnecessary steps, retries or corrections.
Why total task time is more useful than token price alone
For long-running agents, cost per million tokens is an incomplete optimization target. A production team ultimately pays for completed outcomes: model tokens, tool calls, elapsed execution, retries, failed runs and human review. A cheaper model that takes a longer path can lose its list-price advantage. Conversely, a premium model is only economically better when the reduction in corrections, execution time or human intervention is large enough to justify its higher unit cost.
Basis adjusts reasoning effort during the workflow
Basis says its agents increase reasoning effort when a step is difficult and reduce it when the work is easier. OpenAI's customer story says the model can make those changes while keeping its cache intact. The pattern is important because long tasks are rarely equally difficult from beginning to end: extraction or formatting may need little reasoning, while an ambiguous accounting decision may deserve substantially more.
The internal eval measures process, not only the final answer
Basis says its evaluation checks how the agent works as well as what it produces. Examples include whether it follows required templates, consults primary sources for tax questions and checks its own work. The company reports about a 20% improvement in these internal evaluation scores with Astra and attributes part of that gain to better understanding of user intent.
This is a useful design principle even if another team never uses Astra. Agent evaluations should include trajectory requirements that matter to the business. A correct-looking final document is not enough if the agent skipped a required source, ignored an approval step or created hidden rework for a human reviewer.
What the public case study does not prove
- It does not establish that Astra will cut every long-running agent task by 50%.
- The public material does not disclose a full sample size or repeated-run distribution for the workbook comparison.
- It does not publish enough accuracy and correction data to calculate a verified cost per correct completed workbook.
- The ~20% result refers to Basis's internal evaluation, so it should not be compared directly with a public benchmark score.
- Results from a structured accounting workflow may not transfer to research, coding, sales or marketing agents.
A better scorecard for long-running agents
Measure the completed job
| Metric | What it tells you |
|---|---|
| Successful completion rate | How often the agent reaches the required outcome without a restart |
| Total task cost | Tokens, tool calls and other execution costs for a completed job |
| Elapsed task time | Whether the workflow is operationally fast enough |
| Human correction time | How much hidden labor remains after the agent finishes |
| Retry / recovery rate | How often the agent takes a wrong path or needs another attempt |
| Process compliance | Whether required sources, templates, approvals and checks were actually used |
How to test a model upgrade in your own workflow
- Choose a real recurring task with a clear definition of successful completion.
- Run the current and candidate models on the same representative task set.
- Record total execution time, tokens, tool calls, retries and human corrections.
- Score process compliance separately from final-output quality.
- Calculate cost per successful completed task rather than comparing token price alone.
- Route only the workloads that show a meaningful measured benefit to the more expensive model.
Basis's reported 50-tab workbook result is interesting evidence, not a universal benchmark. The reusable lesson is stronger: evaluate agents as end-to-end workers. Measure whether they finish correctly, how directly they get there, what they cost after retries and review, and whether they follow the process your business actually requires.
Sources & useful resources
- OpenAI: Basis completes a tax workbook 2x faster with GPT-6 Astra— Primary vendor customer story, Sep. 28, 2026
- OpenAI: GPT-6 Astra— Official model capabilities and positioning