OpenAI has released GPT‑6.1 Sol, the successor to GPT‑6 Sol, in Codex, ChatGPT Work and its API. The company says the upgrade brings performance closer to GPT‑6 Astra in agentic programming, computer use and professional work. Its reported improvements cover software engineering, document questions and the completion of multistep business processes.
Coding tasks and complex PDF questions
For software development, OpenAI reports that GPT‑6.1 Sol matched GPT‑6 Astra in DeepSWE v1.1, which evaluates complex software-engineering tasks in real repositories. The new model also exceeded the best result achieved by GPT‑6 Sol by 6.4 percentage points. That improvement came at a lower reasoning-effort setting than the one used for its predecessor’s best score.
In document work, OpenAI says GPT‑6.1 Sol scored above Opus 5.5 with fallbacks in GDP.pdf across the tested reasoning-effort settings, while approaching Astra’s performance. The benchmark measures how accurately models answer professional questions using complex PDF documents. Its questions draw on material that includes tables, charts and diagrams, as well as details in fine print.
Business workflows and computer use
OpenAI reports a 2.2-percentage-point lead over Opus 5.5 in AutomationBench at medium reasoning effort. At that same setting, GPT‑6.1 Sol improved on GPT‑6 Sol by 4.8 percentage points. AutomationBench 1.0.6 assesses whether agents correctly complete business workflows from beginning to end, with multiple steps and access to 47 tools across sales, marketing, operations, support, finance and HR.
The computer-use results come from the offline set of OSWorld 2.0, which evaluates demanding workflows spanning everyday and professional tasks. At maximum reasoning effort, OpenAI reports that GPT‑6.1 Sol exceeded GPT‑6 Sol by seven percentage points. It remained 2.1 points behind GPT‑6 Astra, also tested at maximum effort. The reported measure is partial reward on the offline set from release v2026.08.08.
Scientific workflows improve, with Astra still ahead
For scientific work, OpenAI says GPT‑6.1 Sol more than doubled GPT‑6 Sol’s score in Terminal-Bench Science 0.1 when tested at maximum reasoning effort. This evaluation covers scientific workflows including data analysis, simulation and theorem proving. Despite the improvement over the previous Sol version, GPT‑6 Astra retained the highest score among the models tested, at 68.1%, according to the company.
Factual errors and reporting broken tools
OpenAI measured factual accuracy using de-identified conversations in which users had flagged an earlier model’s error. At low reasoning effort, the share of answers containing at least one factual error fell from 11.4% for GPT‑6 Sol to 7.7% for GPT‑6.1 Sol. These figures apply to a deliberately difficult selection of conversations, rather than a sample representative of ordinary use.
A separate safety evaluation tested whether agents informed users when their search tool was broken, rather than responding with a best guess. With reasoning effort set to maximum, OpenAI recorded a failure to disclose the problem in 2.1% of cases for GPT‑6.1 Sol. The corresponding figures were 4.9% for GPT‑6 Sol and 1.5% for GPT‑6 Astra. The tasks were selected to provoke failures, not to measure their frequency in typical use.
Access and API pricing
GPT‑6.1 Sol is available in ChatGPT Work and Codex to users on the Plus, Pro, Business, Enterprise and Edu plans. The model is not yet available in Chat. Developers have a separate access route through the OpenAI API, where the released model uses the identifier gpt-6.1-sol.
Standard API rates are $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. OpenAI says the standard input and output rates are one-fifth of Astra’s. According to the company, cached input is 95% cheaper than standard input and 50% cheaper than cached input for GPT‑6 Sol. That cached rate benefits agents that reuse the same context across requests.
A faster Codex variant is planned
OpenAI plans to offer GPT‑6.1 Sol Ultrafast in the coming days. The company says the planned variant will generate tokens in Codex up to eight times faster than the standard version, with the speed comparison specifically applying to token generation.



