StackAI browser agent cuts model costs 76x using GPT‑6.1 Sol

Asana dramatically reduced both the cost and latency of its StackAI browser agent by using GPT‑6 Astra in Codex to analyze, refactor and experimentally optimize its browser automation workflows. In a controlled 144‑run study, the company tested GPT‑6.1 Sol alongside three other frontier models and found a configuration that brought estimated model cost per run down to $0.47 and cut average runtime to roughly four minutes.

The work targeted StackAI, Asana’s acquired platform for no‑code cross‑system automation that runs browser agents to navigate websites, fill forms and collect data. Small inefficiencies that are negligible at one run can multiply into large cost and latency penalties at scale; to address that, StackAI CTO Frank Hidalgo directed GPT‑6 Astra in Codex to map the agent’s behavior, propose fixes and execute experiments. Hidalgo estimated the same effort would have taken one to two months to do manually; with GPT‑6 Astra it took about a week.

GPT‑6 Astra began by producing a detailed mapping of the agent’s codebase and explaining how model requests were composed. The analysis revealed a key inefficiency: while the agent cached static instructions and tool definitions, it did not cache the growing history of page text and screenshots that the agent collected during navigation. As a result, each model request retransmitted the accumulating history at full cost. The agent’s frequent edits to stored history — dropping older screenshots and repeatedly trimming page text — also meant a naive static cache would not have prevented repeated retransmission or the need to revisit already‑read pages.

Hidalgo selected three recommended changes to test: extend caching to include browsing history, increase the retained text budget, and change screenshot handling so screenshots would be removed in batched intervals instead of at every step. Because the codebase was not set up for controlled trials, Astra refactored portions of the system so multiple workflows could run in parallel under distinct settings. The team then executed a full factorial study across histories, caching and screenshot policies.

The experimental design ran each configuration three times across four models — labelled in the study as Model A (a smaller frontier lab model), Model B (the original production model), Model C (an updated variant of Model B) and GPT‑6.1 Sol. Two history budgets were tested, 120,000 and 480,000 characters, and six different caching and screenshot policies. Each run executed the same representative task: collecting six fields for each of 32 books from a public demo catalog.

The best performing policy combined a larger 480,000‑character history budget with a screenshot policy that allowed screenshots to accumulate to 20 before pruning back to the most recent one. This approach preserved more earlier context between removals and kept a larger share of input eligible for cache reuse. In the optimized runs on GPT‑6.1 Sol, 89% of input came from cache, which was priced at roughly 5% of uncached input cost — a major factor in driving per‑call savings.

Results were substantial. For Model B, the optimizations reduced estimated model cost per run from at least $36.21 (some baseline runs hit a step limit and did not finish) to $1.24, a 29x reduction. Running the same optimized workflow on GPT‑6.1 Sol lowered average cost to $0.47 per run — a 76x reduction relative to the original production configuration on Model B. Runtime also improved: a process that previously took at least 22.5 minutes on Model B was shortened to about four minutes on GPT‑6.1 Sol, a roughly fivefold speedup. The larger history budget additionally improved reliability: on GPT‑6.1 Sol the bigger budget increased completed runs producing correct answers from three of 18 to all 18 in the study.

All experiment sessions, requests, traces and outputs were recorded in Asana’s Command platform. The engineering team converted findings from Command into tickets and pull requests and deployed the changes to production. Asana has released the browser navigation changes in StackAI and is building tooling to make similar experiments easier to repeat, with plans to incorporate these tests into platform evaluations so customers and internal teams can compare cost, runtime and answer quality when configuring agents.

Beyond this cost optimization effort, Asana is also using GPT‑6 Astra in Codex to exercise product features before release: Astra navigates the platform, exercises inputs and surfaces bugs for human QA reviewers. Hidalgo characterized the work as foundational to a development lifecycle in which many cloud agent sessions run in parallel to validate features.

The study demonstrates how targeted cache and history management, combined with automated analysis and controlled experimentation powered by advanced models, can produce outsized reductions in model spend and latency for browser agents. Asana’s published findings and the StackAI updates provide a practical template for teams seeking to lower the operational costs of large‑scale agent deployments.

Source: Read the original source

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *