OpenAI’s GPT-6.1 Sol gets closer to Astra at a fraction of the cost

OpenAI’s GPT-6.1 Sol gets closer to Astra at a fraction of the cost

New Delhi: OpenAI has released GPT-6.1 Sol, an upgrade to GPT-6 Sol, and the pitch is simple: nearly GPT-6 Astra’s ability without Astra’s bill. The company says the new model comes close to Astra on agentic coding, computer use and professional work.

Most of the numbers, all from OpenAI’s own testing, lean on cost. On DeepSWE v1.1, which tests complex software tasks in real codebases, the model matches Astra at roughly a fifth of the price. It also beats GPT-6 Sol by 6.4 percentage points while using less reasoning effort.

Document work follows the same pattern. On GDP.pdf, built around PDFs packed with tables, charts and fine print, it outscores Opus 5.5 with fallbacks at under half the cost per task. It gets near Astra at about a fifth. On AutomationBench, which checks whether agents finish multi-step business workflows, it sits 2.2 points above Opus 5.5 at medium effort for about a third of the cost. On the offline set of OSWorld 2.0, a computer-use test, it’s seven points ahead of GPT-6 Sol at maximum effort and within 2.1 points of Astra, at roughly a seventh of Astra’s cost per task.

The science results show the biggest price gap. On Terminal-Bench Science 0.1, at maximum effort, an average task costs $5.47 with 6.1 Sol. That compares with $23.21 for Opus 5.5 and $23.80 for Astra. Sol more than doubles GPT-6 Sol’s score there. Astra still posts the best result at 68.1%, though, and OpenAI itself says Astra is the one to use for the hardest research.

Factuality improved too. The share of responses containing a factual error falls from 11.4% to 7.7% at low effort, about 32% fewer. But the test prompts came from conversations where users had already flagged an earlier model’s mistake, so they’re deliberately hard. OpenAI says they don’t reflect typical use.

That caveat applies more broadly. These are OpenAI’s benchmarks, and until independent testers weigh in, treat the comparisons with some care.

On safety, OpenAI reports lower failure rates than GPT-6 Sol in three areas: being upfront about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. It saw no attempts to bypass an automated safety reviewer, the same as Astra and GPT-6 Sol. Those evaluations are also built around difficult situations, not everyday use.

The model is available starting today to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It isn’t in Chat yet. Developers can call it as gpt-6.1-sol through the API at $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. The cached rate is 95% below standard input pricing and half of what GPT-6 Sol charged, which should help anyone running agents that reuse context.

OpenAI also plans to release GPT-6.1 Sol Ultrafast in the coming days, with up to eight times faster token generation in Codex.

Punit Panchal
Senior Editor

I’m a content writer specializing in tech, creating clear, engaging, and SEO-friendly content that simplifies complex topics. From emerging technologies to product insights, I focus on delivering value-driven content that connects with readers and ranks effectively.

Comments are closed