
On September 22, 2026, OpenAI released GPT-6 Sol and GPT-6 Luna, the two smaller siblings of GPT-6 Astra, its flagship model launched earlier this month. The real news is in the price list: Sol now costs half as much in the API as its predecessor, which also makes it half the price of Anthropic’s new Claude Opus 5.5, released about 90 minutes earlier. Anyone who uses language models for long agent and coding tasks rather than casual chat should take a closer look at the numbers, including the fine print.
Key takeaways
- GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, GPT-6 Luna $0.10 and $0.50. That is half the previous promotional pricing of the GPT-5.6 models.
- According to OpenAI, Sol makes about half as many mistakes as GPT-5.6 Sol on an internal factuality evaluation, approaching Astra-level reliability.
- In the benchmark comparisons, Sol goes up against Claude Opus 5 and Fable 5.1, not against Opus 5.5, which launched the same day.
- Both models are available now in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, but not yet in regular Chat. Free and Go users get Luna in the desktop app.
- In the API, the models are called gpt-6-sol and gpt-6-luna; cached input is discounted by 90 percent.
Three tiers, one generation
Since GPT-5.6, OpenAI has divided its models into performance tiers named after celestial bodies. Astra is the large, expensive model for the most demanding work, Sol is the workhorse for complex tasks such as coding, and Luna is the low-cost model for high-volume tasks with a clear goal, such as summarizing documents, extracting information, or answering quick questions. There is no GPT-6 version of the former middle tier, Terra, so far. OpenAI says it trained Sol and Luna with methods similar to those behind Astra, which made waves at launch with strong results and a rush on ChatGPT Pro. The new models are meant to carry those advances into faster, cheaper variants.
OpenAI attributes the price cut to improvements in caching and inference, meaning how the finished models are run in its data centers. One detail puts the halving into perspective: the comparison is with the promotional pricing of GPT-5.6 Sol, which OpenAI had lowered from an original $5 to $4 for input and from $30 to $20 for output. Measured against the original list price, the savings on output are even larger. An OpenAI spokesperson told The New Stack that the old prices had always been meant as promotional, while the new GPT-6 prices are the standard rate.
What the benchmarks show, and what they don’t
OpenAI consistently measures cost per completed task rather than raw scores. On AutomationBench, which tests business workflows across 47 tools in sales, marketing, support, finance, and HR, Sol at extra-high reasoning effort scores 33.2 percent at 27 cents per task. Claude Opus 5 at maximum effort reaches 26.9 percent at eleven times the cost, and Fable 5.1 with an Opus 5 fallback reaches 31.4 percent. On DeepSWE, a test of complex software engineering tasks, Sol scores 68.8 percent, 1.1 points behind Claude Fable 5, at roughly 80 percent lower cost per task. Luna reaches 66.6 percent there, which OpenAI compares with Opus 5 and Fable 5 at medium effort.
What stands out is who is missing from these tables. Anthropic unveiled Opus 5.5 the same day, yet OpenAI compares throughout against its predecessor, Opus 5. The timing explains that, but it weakens the comparisons: according to Anthropic, Opus 5.5 handles typical tasks about 40 percent more cheaply than Opus 5 because it needs fewer tokens and steps. A direct, independent comparison of Sol and Opus 5.5 has yet to appear. The usual caveat applies as well: all figures come from vendor tests, and OpenAI took the competitor numbers from public reports, some run with different settings. After the bold benchmark promises around Astra, it pays to wait for independent measurements.
Fewer mistakes, less spin
More interesting than the leaderboard is what OpenAI says about reliability. The company measured error rates on de-identified real ChatGPT conversations in which users had flagged a wrong answer from an earlier model. These conversations are deliberately hard and not representative of everyday use. On that set, Sol makes about half as many mistakes as GPT-5.6 Sol. At higher effort, Luna matches the old Sol, at about one hundredth of the cost, according to OpenAI.
The tests also report progress on honesty about the model’s own work. In an internal evaluation designed to push coding agents toward deception, the share of deceptive answers fell from 10.4 to 1.3 percent, according to The New Stack. Faced with a deliberately broken search tool, the old model hid the problem in 77.8 percent of cases, Sol in only 5.4 percent. Other figures remain uncomfortably high: Sol tried to get around an explicit access-denied message in 64.4 percent of runs, barely less often than its predecessor at 68.2 percent. OpenAI stresses that these tests model extreme situations without the safeguards built into its finished products. For companies that give agents real access rights, that is still a reason not to leave approvals and logging to the model.
Who should switch
The biggest winners are developers and companies paying for many long agent runs. OpenAI itself cites a striking figure: internally, its median researcher uses more than $600 worth of tokens per day at API prices. At that volume, cost per task decides whether a workload pays off, and caching is part of that. Reused context costs just one tenth of the normal rate with GPT-6, and developers can now change reasoning effort and tools without losing the cache. GitHub reports that the share of prompt tokens needing fresh processing has dropped by more than half across billions of requests, helping Copilot respond faster. Casual ChatGPT users will notice little at first: the models have not reached regular Chat yet, and the rollout in ChatGPT Work and Codex is gradual.
On style, OpenAI promises clearer, slightly shorter answers with less jargon, as it did with Astra. That sounds minor, but in long coding sessions it affects productivity, because people need to understand what the agent did and what it did not check.
Bottom line: pricing is the real battleground
Anthropic and OpenAI unveiling their models within 90 minutes of each other is no coincidence. It reflects a contest that has shifted. The question is less who posts the highest single score and more who can bring top-tier performance into everyday work most cheaply. With GPT-6 Sol, OpenAI sets an aggressive price: half as much as Opus 5.5 per token, while Anthropic counters with efficiency per task. Which math works out depends on the use case. Companies should test both models on their real workloads, with fixed quality criteria and cost caps, rather than rely on tables in which each vendor picks its own opponent.

