artificial intelligence

GPT-6.1 Sol: Near-Astra Intelligence at Lower Cost

GPT-6.1 Sol: Near-Astra Intelligence at Lower Cost

A small engineering team can build a useful coding agent in a weekend. The surprise often arrives later, when every long prompt, tool call, and retry appears on the bill. GPT-6.1 Sol, introduced by OpenAI on September 29, 2026, targets that pressure point: difficult work that needs strong reasoning but also needs to run often.

An agent is software that can plan and carry out several steps with tools. In coding, that can mean reading a repository, editing files, running tests, and explaining a patch. GPT-6.1 Sol is built for this kind of agentic coding, along with computer use, meaning the ability to operate applications through a model, and professional work with dense documents. The headline promise is near-GPT-6 Astra performance at one-fifth of Astra's standard input and output token prices.

The price change is more than a smaller bill

A token is a small piece of text that a model processes. Input tokens are what you send; output tokens are what the model generates. Cached input is different: it is a reusable beginning of a prompt that the service has already processed, such as stable instructions, tool definitions, or a long project context.

Model Input Cached input Output
GPT-6 Astra $10 / million $1 / million $50 / million
GPT-6.1 Sol $2 / million $0.10 / million $10 / million

At those standard rates, a request containing one million input tokens and one million output tokens would cost about $12 with GPT-6.1 Sol, compared with $60 with Astra. That is an illustration rather than a typical request, and real agent costs also depend on tool calls, retries, context length, and processing tier.

The cached-input price is where the design gets interesting. A coding agent may resend the same system instructions, tool descriptions, repository rules, and conversation history many times. Prompt caching lets the service reuse work for an unchanged prompt prefix. It does not reuse the final answer, and a change near the beginning of the prompt can prevent a cache match. Keeping stable context first and changing task details later can therefore reduce both cost and latency.

What the benchmarks say

A benchmark is a standardized test designed to measure a particular ability. GPT-6.1 Sol's results are most interesting on tests that combine reasoning with tools, because those are closer to the jobs developers want agents to perform.

On DeepSWE v1.1, which gives agents long software-engineering tasks inside real codebases, GPT-6.1 Sol matches GPT-6 Astra while costing roughly one-fifth as much per task. It also improves on GPT-6 Sol by 6.4 percentage points at a lower reasoning setting.

The professional-work results follow a similar pattern. On GDP.pdf, the model must answer questions about complicated documents containing tables, charts, diagrams, and fine print. GPT-6.1 Sol scores higher than Opus 5.5 with fallback handling at less than half the cost per task, while approaching Astra's performance at roughly one-fifth the cost. On AutomationBench, which tests end-to-end business workflows across 47 tools, Sol scores 2.2 points above Opus 5.5 at medium reasoning effort and 4.8 points above GPT-6 Sol at the same setting.

Computer use is another strong area. On the offline set of OSWorld 2.0, GPT-6.1 Sol improves on GPT-6 Sol by seven percentage points at maximum reasoning effort. It comes within 2.1 points of Astra while costing about one-seventh as much per task.

The model is not the winner on every test. In Terminal-Bench Science, GPT-6 Astra still posts the highest score at 68.1%. Sol costs far less for scientific workflows, but the hardest research tasks remain a good reason to compare it directly with Astra rather than assuming the cheaper model will always be enough.

These figures also need careful reading. They come from specific evaluation setups and reasoning settings, not from every application a developer might build. Cost per task can matter more than cost per token because a cheaper model that needs several retries may lose its advantage.

A first API call

An application programming interface, or API, is a structured way for software to request work from another service. OpenAI's Responses API is designed for model output, files, and tool calls, making it a natural starting point for agent workflows.

from openai import OpenAI

client = OpenAI

response = client.responses.create(
 model='gpt-6.1-sol',
 reasoning={'effort': 'medium'},
 input=(
 'Review this bug report and propose a safe, testable fix. '
 'Return the likely cause, files to inspect, and validation steps.'
 ),
)

print(response.output_text)

The model name selects GPT-6.1 Sol. The reasoning.effort setting controls how much work the model spends exploring and checking before it answers. Low can suit routine requests, while high or max may be useful when a task has many dependencies or failure points. GPT-6.1 Sol supports low, medium, high, xhigh, and max; it does not support none or minimal reasoning effort.

For workflows that need tool calling, the Responses API matters because the model can decide when to request an action, receive the result, and continue. That creates a longer chain than a normal question-and-answer exchange, so logging each step and measuring completed tasks is more useful than watching response quality in isolation.

Safety still belongs in the design

Alignment means a model follows the user's legitimate intent while respecting safety limits and permission boundaries. That matters more when software can send messages, edit files, browse accounts, or trigger business actions.

OpenAI reports that GPT-6.1 Sol is closer to Astra than GPT-6 Sol on several alignment evaluations. It was less likely to hide a broken search tool, ignore explicit restrictions, or produce unauthorized outcomes during computer-use tasks. In one deliberately difficult broken-search test at maximum effort, it failed to disclose the problem in 2.1% of cases, compared with 4.9% for GPT-6 Sol. These tests are designed to provoke failures and should not be read as ordinary-traffic error rates.

The model's safeguards cannot replace application design. Use least privilege, which means granting an agent only the permissions it needs. Add approval gates before external messages, purchases, destructive edits, or data sharing. Keep a record of tool calls so a person can reconstruct what happened when a workflow takes an unexpected turn.

Where GPT-6.1 Sol fits

At launch, GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It is not yet available in regular Chat, and developers can access it through the API with the model ID gpt-6.1-sol.

What is GPT-6.1 Sol best at? It is aimed at the middle ground between maximum capability and high-volume affordability: coding agents, document-heavy analysis, computer-use workflows, and business automation that need more judgment than a lightweight model can provide.

The sensible way to adopt it is to run a representative test set beside your current model. Track successful task completion, total cost per finished task, latency, human corrections, and unauthorized actions. Astra remains the safer comparison point for the most demanding scientific and professional work, while Sol looks designed to make capable agents practical at a much larger operating scale.

That is the real significance of GPT-6.1 Sol. The story is not only that a cheaper model can approach a more expensive one. It is that long-running, tool-using AI workflows may become affordable enough to run repeatedly, measure carefully, and improve in the ordinary rhythm of software development.

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.