Basis halves tax workbook time using GPT-6 Astra

Basis, an AI firm focused on automating routine accounting tasks, reported that GPT-6 Astra completed a complex 50-tab tax workbook in roughly half the time required by the company’s prior model, GPT-5.6 Sol. According to Basis co-founder Mitch Troyanovsky, the faster completion stemmed from the model making stronger early decisions and following a more direct execution path, which reduced rework and token consumption over long runs.

The evaluation compared end-to-end agent performance on a multi-sheet tax workbook: filling, validating, and producing final tax outputs across 50 tabs. Basis observed that GPT-6 Astra moved through the spreadsheet with fewer corrections and less backtracking than GPT-5.6 Sol. That difference in approach translated to a sizeable reduction in total runtime for the same task, an outcome Basis attributes to improved initial choices and more efficient planning by the newer model.

Beyond raw speed, Basis highlighted two technical behaviors in GPT-6 Astra that mattered for long-running workflows. First, the model demonstrated an ability to vary its reasoning effort as a task progressed: it could allocate more computational intensity to challenging steps and scale back for straightforward ones. Second, Astra maintained a cache of useful context across steps, preserving information that helped subsequent decisions without repeated recomputation. Together, these features lowered unnecessary computation and reduced token usage during extended processes.

Those capabilities have practical implications for businesses that automate extensive accounting workflows. By matching reasoning effort to subtask difficulty, agents powered by GPT-6 Astra can avoid the cost of uniformly heavy computation while still applying deeper reasoning where it’s needed. Basis said this adaptive behavior helps shorten response times for routine portions of a workflow and reduce overall operational expense for longer automation runs.

Reliability and intent recognition were additional areas where Basis reported gains. In internal evaluations, the company measured roughly a 20% improvement in scores when agents used GPT-6 Astra versus GPT-5.6 Sol. Basis credits this uplift to the model’s stronger understanding of user goals and broader contextual cues. That understanding helped agents decide when to ask clarifying questions, when to flag assumptions, and when to adhere to explicit instructions—reducing the amount of bespoke rule-writing needed to handle edge cases.

Reducing the need for detailed case-by-case rules matters for real-world deployment. Basis argued that better intent recognition enables agents to generalize to scenarios that differ from internally tested examples, increasing confidence that automation will behave correctly when unexpected variations arise. For accounting teams, that translates into fewer maintenance burdens and a lower likelihood of manual intervention during production runs.

The improvements Basis identified also align with the company’s broader mission of shifting accountants away from repetitive manual tasks toward higher-value strategic work. Faster, more reliable agents can shorten project turnaround times and cut the manual effort required to validate complex spreadsheets and tax documents. Basis emphasized two practical advantages: higher throughput on multi-step tasks and a reduced need for bespoke guidance to steer agent behavior. For organizations that regularly rely on multi-sheet workbooks and long-running processes, those advantages can mean lower operational costs and less maintenance overhead.

Basis published its evaluation through OpenAI, and the company’s results underscore how advances in base models can translate into tangible productivity improvements for domain-specific automation. While the report focused on one representative tax workbook task, the findings suggest that adaptive reasoning, preserved context across steps, and stronger intent recognition can materially improve agent effectiveness for other extended workflows as well.

In summary, Basis’s tests found that GPT-6 Astra halved completion time on a 50-tab tax workbook compared with GPT-5.6 Sol, improved agent decision-making, and raised internal evaluation scores by about 20%. These changes reduced rework, lowered token consumption during long runs, and lessened the need for detailed rule-writing—outcomes that could make accounting automation faster, more economical, and easier to deploy in production. Basis’s report, published by OpenAI, provides the company’s assessment of Astra’s impact on its workflow automation efforts.

Source: Read the original source

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *