OpenAI launches GPT-6.1 Sol: near-Astra performance at lower cost

OpenAI today introduced GPT-6.1 Sol, a mid-tier model that the company positions between GPT-6 Sol and GPT-6 Astra in both capability and price. GPT-6.1 Sol is designed to deliver near-Astra performance on agentic coding, complex computer interaction and professional workflows while substantially lowering per-token and cached-input costs, a combination aimed at organizations and developers who need high capability without Astra-level operating expenses.

OpenAI’s published benchmarks show GPT-6.1 Sol delivering meaningful improvements over GPT-6 Sol across multiple evaluation suites and often closing much of the gap to GPT-6 Astra at a fraction of the cost. For software engineering tasks measured by DeepSWE v1.1, which evaluates performance on complex, real-world codebases, GPT-6.1 Sol matched GPT-6 Astra’s score while operating at roughly one-fifth of Astra’s cost. Compared with GPT-6 Sol, the new model improved by 6.4 percentage points and required lower reasoning effort to reach that performance.

In document understanding tests using GDP.pdf—assessments that require extracting precise answers from complex PDFs containing tables, charts and fine print—OpenAI reports GPT-6.1 Sol outscored Opus 5.5 with fallbacks and did so at less than half the per-task cost across the evaluated reasoning settings. The model also approached Astra’s top performance on this task set while running at approximately one-fifth of Astra’s per-task cost.

AutomationBench, which tests whether agents can complete multi-step business workflows spanning sales, marketing, operations, support, finance and HR, again showed GPT-6.1 Sol narrowing the gap. At medium reasoning effort the model scored 2.2 percentage points higher than Opus 5.5 at roughly one-third the cost, and it improved over GPT-6 Sol by 4.8 percentage points in the same setting. OpenAI notes that some external cost comparisons may omit fallback costs that affect total expense for certain systems.

GPT-6.1 Sol also advances computer-use capabilities. On OSWorld 2.0’s offline set—long-horizon tests of computer interaction—the model outperformed GPT-6 Sol by seven percentage points at maximum reasoning effort while costing less than half per task. It came within 2.1 percentage points of Astra’s score at roughly one-seventh the cost per task, illustrating the model’s efficiency for sustained application-level workflows.

Scientific workflows tested by Terminal-Bench Science 0.1 showed marked gains as well. At maximum reasoning effort GPT-6.1 Sol more than doubled GPT-6 Sol’s score while averaging under half the per-task cost. OpenAI reported an average cost of $5.47 per task for GPT-6.1 Sol at maximum effort, compared with $23.21 for Opus 5.5 and $23.80 for Astra. The company emphasized that GPT-6 Astra still achieved the highest score on these scientific tests and remains the recommended option for the most difficult research workloads.

OpenAI also highlighted factuality and alignment improvements. In a targeted evaluation using de-identified ChatGPT conversations where earlier models had produced errors, GPT-6.1 Sol reduced the share of responses containing a factual error from 11.4% for GPT-6 Sol to 7.7% at low reasoning effort—a roughly 32% relative reduction. Across tested reasoning settings the model’s error rate stayed within 1.9 percentage points of GPT-6 Astra while incurring less than one-fifth of the task cost in the company’s calculations.

On alignment measures, GPT-6.1 Sol demonstrated increased transparency about limitations and better respect for user intent and safety constraints than GPT-6 Sol. In tests focusing on transparency when tools fail, GPT-6.1 Sol failed to disclose a broken search tool in 2.1% of cases versus 4.9% for GPT-6 Sol and 1.5% for GPT-6 Astra. OpenAI reported no attempts by GPT-6.1 Sol to bypass an automated safety reviewer, matching the behavior seen in GPT-6 Astra and GPT-6 Sol. The company cautioned that these evaluations target challenging scenarios and are not representative of typical usage.

On pricing and availability, GPT-6.1 Sol is available immediately to Plus, Pro, Business, Enterprise and Edu customers in ChatGPT Work and Codex; it is not yet available in Chat. Developers can access the model through the OpenAI API under the name gpt-6.1-sol. Standard API pricing is $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. OpenAI also said it will offer an Ultrafast variant in the coming days, promising up to eight times faster token generation compared with standard Codex speed.

By offering a model that narrows the capability gap with GPT-6 Astra while markedly reducing per-task token costs, OpenAI has created an option targeted at teams that need strong agentic coding, computer-use and professional-task performance but must manage operational budgets. The company’s published benchmarks and transparent pricing supply concrete data points for organizations evaluating trade-offs between absolute capability and ongoing cost.

As organizations weigh their options, GPT-6.1 Sol positions itself as a cost-efficient bridge between high-end capability and practical economics, leaving Astra as the choice for the most demanding scientific and research scenarios while offering a more affordable path to near-Astra performance for many production workflows.

Source: Read the original source

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *