GPT-6 Astra got the headlines when it launched earlier this month, but the more interesting move might be what OpenAI just did below it. The company announced GPT-6 Sol and GPT-6 Luna, two models positioned beneath Astra in the family hierarchy but trained using similar methods. And they’re 50% cheaper than the GPT-5.6 models they replace. That kind of price drop, paired with genuine benchmark improvements, is worth paying attention to.
The pitch here is straightforward: not every task needs Astra. Research agents running overnight, coding pipelines processing millions of tokens daily, business workflow automation at scale — these are cost-sensitive workloads. OpenAI is betting that Sol and Luna can handle most of that load while keeping bills manageable.
What you’re actually paying now
The new API pricing is blunt and easy to read. GPT-6 Sol drops from $4 to $2 per million input tokens, and from $20 to $10 on output. GPT-6 Luna goes from $0.20 to $0.10 on input and $1.20 to $0.50 on output. OpenAI says the cuts come from improvements in caching and inference efficiency, and that the savings are being passed directly to customers rather than held as margin.
For context, GPT-6 Astra still sits at the top and is the right call when you need the best result regardless of cost. But Sol and Luna are meant to cover the wide middle ground where most production workloads actually live.
Benchmark performance against Anthropic’s lineup
The competitive framing here is clearly aimed at Anthropic. On AutomationBench, a test of business workflows across 47 tools spanning sales, finance, HR, and operations, GPT-6 Sol at extended effort outperforms Claude Opus 5 at max effort — and does it at roughly 9% of Opus 5’s cost per task. That’s not a small gap. Sol also beats Claude Fable 5.1 with Opus 5 fallback, which is notable because that Fable 5.1 figure reportedly understates its true cost since fallback calls to Opus 5 occurred on around 40% of tasks.
On Agents’ Last Exam, which evaluates long-horizon professional tasks across 55 sub-industries, Sol at max effort scores 56.4% versus Opus 5’s highest result — at 60% lower cost per task. These are the kinds of benchmarks that matter to teams building serious agent systems, not just running one-off queries.
Coding performance is also a key part of the story. On DeepSWE, a test of real-world software engineering tasks in actual codebases, Sol at max effort scores 68.8%, within about one percentage point of Claude Fable 5’s best result. Luna scores 66.6%, comparable to Opus 5 and Fable 5 at medium effort, at 93% and 96% lower cost respectively. OpenAI notes that internal researcher token usage has grown sharply, with median daily usage valued at $600 and 90th percentile usage reaching $7,000. Cheaper models matter when usage is that intensive.
Where Sol and Luna actually improve
Beyond benchmarks, a few capability areas stand out:
- Factuality: Sol makes roughly half as many factual errors as GPT-5.6 Sol, based on real-world conversations where users flagged mistakes. Luna at higher effort levels matches GPT-5.6 Sol at a fraction of the cost.
- Coding: Sol matches Claude Fable 5.1 on FrontierCode, which grades code on mergeability, not just correctness. That includes test quality, code style, and codebase adherence.
- Computer use: Astra still leads here, but Sol and Luna improve on their predecessors’ cost efficiency for GUI and system interaction tasks.
- Caching: Improved caching is specifically designed to benefit agent workflows and long conversations, where repeated context means more potential savings.
Why this matters beyond the price tag
The race between OpenAI and Anthropic is increasingly playing out on the cost-intelligence curve, not just raw capability. Google is doing the same thing with its Gemini tiers. What’s shifting is that “frontier” performance is no longer exclusive to the most expensive models. Sol competing with Opus 5 at 9% of the cost isn’t just a marketing stat — it changes which workloads are economically viable to automate.
For developers and founders building on top of these APIs, the calculation is simple. If your use case doesn’t strictly require Astra, Sol and Luna are now significantly more attractive than the previous generation. And with alignment training carrying over from Astra’s methods, there’s less reason to treat the cheaper models as a quality compromise. The GPT-6 family is shaping up to cover more ground than its predecessor, at prices that make large-scale deployment a real option rather than a budget line item.
GPT-6 Sol and Luna are available now via the OpenAI API.



