OpenAI released GPT-6.1 Sol on September 29, 2026, positioning it for teams that need strong coding and computer-use performance without paying frontier-model prices on every request. According to OpenAI's announcement, Sol matches GPT-6 Astra on DeepSWE v1.1 at one-fifth of Astra's standard API cost. That benchmark result does not promise identical results in every application. The practical question is whether Sol meets your quality threshold on your own tasks.
At standard API rates, GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens. Cached input costs $0.10 per million tokens, while cache writes cost $2.50 per million tokens, according to the model documentation. Those prices make it worth testing for coding agents, long-context analysis, and repetitive workflows with reusable prompts.
What is GPT-6.1 Sol? Architecture, Lineup, and Positioning

GPT-6.1 Sol is available through the API as gpt-6.1-sol. Its documented context window is 1,050,000 tokens, with a 128,000-token maximum output. The model supports reasoning settings from low through max; the default is medium. Its published knowledge cutoff is April 30, 2026. These are model specifications, not a guarantee that a million-token request will be economical or equally accurate across every part of the context.
For comparison, GPT-6 Astra lists $10 per million input tokens and $50 per million output tokens. GPT-6 Luna lists $0.10 input and $0.50 output. All three list a 1,050,000-token context window and 128,000-token maximum output. Price and measured task quality are more useful selection criteria than context size alone.
Model | Standard input / 1M tokens | Standard output / 1M tokens |
|---|---|---|
GPT-6 Astra | $10 | $50 |
GPT-6.1 Sol | $2 | $10 |
GPT-6 Luna | $0.10 | $0.50 |
These are standard token rates from each model's documentation. Tool charges, caching behavior, and rates for very long requests can change the total. For Sol requests above 272,000 input tokens, the model page lists higher long-context rates: twice the input and cache rates and 1.5 times the output rate. Estimate costs from your actual request mix before choosing a default.
DeepSWE v1.1: Closing the Software Engineering Gap

OpenAI reports that GPT-6.1 Sol matches Astra on DeepSWE v1.1 and improves by 6.4 percentage points over GPT-6 Sol on that benchmark. The announcement also reports that Sol lands within 2.1 percentage points of Astra on OSWorld 2.0, a computer-use evaluation. Those relative results do not establish a universal success rate for production repositories or desktop tasks.
Use the results as a reason to run a local evaluation. Choose representative tasks: a bug fix with tests, a multi-file feature, a code review that needs evidence, and a task using your actual browser or desktop environment. For each model, record completed tasks, human corrections, elapsed time, token use, and tool errors. An apparently cheaper model can cost more if it causes repeated attempts or review work.
Sol is available in ChatGPT Work and Codex for eligible Plus, Pro, Business, Enterprise, and Edu plans, as described in the release announcement. API availability uses the gpt-6.1-sol model identifier. Check access in your account before changing an automated workflow.
The 95% Prompt Caching Advantage

Sol's $0.10 cached input rate is 95% below its $2 standard input rate. This percentage compares two unit prices. It does not mean a whole agent run will cost 95% less: output tokens, cache writes, uncached input, tools, and long-context surcharges still count.
Suppose a workflow writes 1 million reusable input tokens to cache, then reads the same material 19 more times, and produces 100,000 output tokens. At the listed standard Sol rates, the illustrative bill is $2.50 for the cache write + $1.90 for cached reads + $1.00 for output = $5.40. If none of the 20 million input tokens were cached, input alone would cost $40, plus $1 output. This illustration assumes every repeated read qualifies for the cached rate and ignores fresh tokens, tools, and long-context pricing. Actual cache eligibility and billing depend on the request pattern.
The biggest opportunity is repeated context: stable system instructions, a repository map, schemas, or a reference document reused across many calls. Monitor cache hit rates and billed token categories rather than assuming the discount will appear automatically. If requests exceed the documented 272,000-input-token threshold, include the long-context rate in your estimate.
How to test Sol in an existing workflow
Start with one bounded workload and a clear pass condition. For a coding agent, that might be issues with deterministic tests plus human review of the patch. For a computer-use agent, it might be completing a defined sequence without unintended edits. Record the same measures for Sol and your current model.
If the workflow calls tools, use the Responses API as directed by the Sol model documentation. Keep approval and rollback rules around actions that modify production systems. Compare both model spend and operational work: retries, review time, and failures matter more than the headline input rate.
For predictable routing, use the least expensive model that passes your quality gate on each step. Luna may suit simple classification or extraction; Sol may suit the main coding or reasoning loop; Astra may be reserved for cases where a measured improvement justifies its cost. These are testable routing hypotheses, not fixed roles imposed by OpenAI.
Frequently asked questions
How much does GPT-6.1 Sol cost?
The published standard API prices are $2 per million input tokens, $10 per million output tokens, $0.10 per million cached input tokens, and $2.50 per million cache-write tokens. Longer inputs can use higher rates.
Is Sol as capable as Astra?
OpenAI reports a match on DeepSWE v1.1 and a 2.1-point gap on OSWorld 2.0. Those benchmark results support evaluation for coding and computer-use work. They cannot establish parity on every task.
Does a 95% cache discount reduce my whole bill by 95%?
No. The 95% figure compares Sol's cached-input unit rate with its standard-input unit rate. Your total also includes cache writes, output, uncached input, tools, and applicable long-context rates.
What is the API model name?
Use gpt-6.1-sol. Check the model documentation for current API support and limits.























