Key Takeaways
- Cache reads cut in half:
gpt-6.1-solkeeps the same $2 input / $10 output asgpt-6-sol, while cache reads drop from $0.20 to $0.10 per 1M tokens, just 5% of the standard input price. Cache-heavy workloads such as multi-turn chat, agents and coding assistants will see a noticeably smaller bill - Close to Astra: 52 on the Artificial Analysis Intelligence Index at max effort, 4 points above GPT-6 Sol and only 1 point below GPT-6 Astra. OpenAI says it matches Astra on DeepSWE v1.1 at roughly one-fifth the price
- Fewer factual errors: on deliberately difficult factuality prompts at low reasoning effort, OpenAI reports the share of responses containing a factual error falling from 11.4% to 7.7%
- Live on APIYI: all four billing items match the provider with no markup; savings come from our top-up promotions, see Top-up Promotions
- One change to note: reasoning effort no longer supports
none; the lowest level islow
Background
On 22 September, OpenAI released GPT-6 Sol and Luna, halving the price of its main tier. One week later, on 29 September at DevDay 2026, OpenAI launched an upgraded Sol: GPT-6.1 Sol. The point of this release is getting close to Astra-level capability at Sol pricing. According to OpenAI, GPT-6.1 Sol approaches GPT-6 Astra on agentic coding, computer use and professional document work, while its standard input and output prices are one-fifth of Astra’s. The other change that hits the bill directly is the cache read price, which is halved from $0.20 to $0.10 per 1M tokens. APIYI has launchedgpt-6.1-sol with input, output, cache read and cache write all priced in line with the provider.
In Detail
Benchmarks
Independent evaluation (Artificial Analysis Intelligence Index)
Per Artificial Analysis, GPT-6.1 Sol is 4 points above GPT-6 Sol, 5 points above GPT-5.6 Sol, and 1 point below GPT-6 Astra.
Vendor-reported
OpenAI also frames it in cost terms: on OSWorld 2.0 it is about 2.1 points behind Astra at roughly one-seventh the cost per task; on AutomationBench at medium effort it beats Claude Opus 5.5 by 2.2 points at roughly one-third the cost.
Independent results are from Artificial Analysis (
artificialanalysis.ai/models/releases/gpt-6-1-sol). Vendor-reported figures come from OpenAI’s launch materials as compiled by Vellum and VentureBeat; parentheses show the reasoning effort used, and they have not been independently reproduced. Data retrieved on 30 September 2026.Key Features
Half-Price Cache Reads
Half the cache read price of GPT-6 Sol, just 5% of standard input. Biggest gains where system prompts, tool definitions or long documents are hit again and again
Near-Astra Capability
Close to the flagship on agentic coding, computer use and professional documents, billed at the Sol tier
Better Factuality
Responses with factual errors fall from 11.4% to 7.7% at low effort; OpenAI says it stays within 1.9 points of Astra across all settings
Ultrafast Option
OpenAI says an Ultrafast tier is coming, with up to about 6x generation speed via the API at 6x the standard price. We will announce it separately once it is available on APIYI
Specifications
Practical Use
Recommended Scenarios
- Coding assistants and agents: system prompts and code context are re-read from cache every turn, so halved cache reads show up directly on the bill
- Long-document analysis: ask repeated questions about the same contract, financial report or codebase, with cached tokens billed at just $0.10
- Mid-difficulty tasks currently on Astra: run a comparison with GPT-6.1 Sol first; capability is close and the price is about one-fifth
Code Examples
Best Practices
- Migrating from GPT-6 Sol is mostly a model-name change:
gpt-6-sol→gpt-6.1-sol. The one thing to check is reasoning effort: code that passesnonemust switch tolow - Prefer Responses for tool calling: OpenAI’s model docs list tool calling under the Responses API, so use
/v1/responsesfor function calling and multi-step agents - Put stable content first: place system prompts, tool definitions and long documents at the start of the message so later turns hit the cache at $0.10
- Stay within 272K input where possible: above that, the whole request moves to a higher pricing tier (see the tier table below)
- Leave room for reasoning: reasoning tokens count toward the output limit, so do not set
max_output_tokenstoo tight
Pricing and Availability
Pricing
Provider list prices (USD per 1M tokens, input up to 272K):
The GPT-6.1 Sol column is also APIYI’s pricing: all four billing items match the provider, with no markup.
Model prices are aligned with the provider and may change with them; the table above is for reference only, and the Model Pricing tab in the top navigation is authoritative: Model Pricing.
Long-Context Tiers
As with the provider, when a single request’s input exceeds 272K tokens, the entire request is billed at the second tier:
Tiers are determined per request, not by cumulative account usage.
What Half-Price Cache Reads Save
Take a typical coding-assistant turn: 100K input tokens, of which 90K hit the cache and 10K are new, plus 2K output tokens.
With input and output prices unchanged, the cache alone makes each such turn about 16% cheaper; the higher the cache share and the shorter the output, the more you save.
Groups and Endpoints
See group differences.
Stack With Top-up Promotions
Model prices match the provider, and the savings come through our top-up promotions: bonus credit on top-ups lowers your effective cost further, and top-ups settle at a fixed 1:7 exchange rate. See Top-up Promotions.Summary and Recommendations
GPT-6.1 Sol is a same-price upgrade with better capability and cheaper caching: input and output prices match GPT-6 Sol, capability comes close to GPT-6 Astra, and cache reads are half price. For cache-heavy coding and agent workloads, that is a cost reduction you will see directly on the bill. Recommendations:- Current
gpt-6-solusers: switch togpt-6.1-soland check whether you passreasoning_effort: none; nothing else needs to change - Using
gpt-6-astrafor mid-difficulty tasks: run your own tasks through both; if capability is close, cost drops to about one-fifth - Cache-heavy workloads: keep a stable prefix at the very start of your messages to get the most out of the halved cache price
Sources: OpenAI model docs
developers.openai.com/api/docs/models/gpt-6.1-sol; Artificial Analysis artificialanalysis.ai/models/releases/gpt-6-1-sol; reporting by TechCrunch, VentureBeat, Vellum and others (29 September 2026). Data retrieved on 30 September 2026. APIYI pricing is subject to live platform data.