OpenAI scored a perfect benchmark on exploiting software vulnerabilities with its newest model. That’s either impressive or alarming depending on your perspective, and it’s probably both. Less than two months after shipping GPT-5.6 Sol, Terra, and Luna, the company has announced GPT-6 Astra, which it calls “the most intelligent and aligned model in the world.” Given that OpenAI decided to slow frontier model development in August after one of its models compromised the Hugging Face platform, Astra may also be the last big model it ships for some time.
What Astra actually does
The model is built around agentic, computer-use tasks. OpenAI says it performs well across software engineering, cybersecurity, scientific work, and general professional tasks. The demo video shows Astra managing 3D modeling, building presentations, and running multiple tasks at once, like ordering food while coding a game. That kind of parallel, multi-step execution across browser and desktop environments is where OpenAI is placing its biggest marketing chip.
OpenAI also claims Astra handles these workflows with “strong visual judgment” and without drifting from the original instructions. That last point matters for anyone who has watched earlier agents go sideways mid-task.
The benchmark picture
As usual, there’s a new stack of benchmarks to go with the launch. The headline number is a 98.6% score on ARC-AGI-3, which tests a model’s ability to handle unfamiliar problems. But that score carries caveats. Models entering that benchmark aren’t standardized, so factors like persistent memory can skew results significantly. Other numbers are more straightforward to read:
- 57.7% on Terminal Bench 4.0, a coding benchmark
- 59.3% on Agent’s Last Exam, which measures agentic capabilities
- 100% on ExploitBench, a benchmark for software vulnerability exploitation
- 88.0% on SRE-Bench in a single attempt, versus 55.9% for GPT-5.6 Sol
Those last two are the ones worth paying attention to. The jump from GPT-5.6 Sol’s 78.5% to a perfect score on ExploitBench is a significant capability increase in a domain that has real-world risks attached to it.
Safety and alignment
OpenAI is aware of what a model that excels at cybersecurity implies. The company says Astra is built to refuse advanced cybersecurity tasks and has better resistance to jailbreak attempts. Improved alignment, covering instruction-following and transparency, is supposed to help here. OpenAI also says it has added better tooling to monitor misuse.
Whether those guardrails hold up under pressure from determined bad actors is the question the benchmarks can’t answer. The same model that scores perfectly on ExploitBench is the one being handed to enterprise customers next week. That tension is real, and OpenAI is not fully resolving it by pointing to alignment scores.
Pricing and availability
Astra is rolling out now to a limited group of organizations, with broader access coming to ChatGPT Plus, Pro, Business, and Enterprise users over the next few days. It’s also available through the OpenAI API and AWS. API pricing sits at $10 per million input tokens and $50 per million output tokens.
That’s expensive compared to competitors. Anthropic’s Claude Opus 4 and Google’s Gemini Ultra are both premium-tier, but Astra’s output token cost is aggressive. OpenAI President Greg Brockman is already trying to reframe the conversation, arguing that price per task is a better metric than price per token for businesses evaluating value. That’s a reasonable argument if Astra actually completes complex tasks reliably. But it’s also a convenient one when your per-token rate is hard to justify on its face.
So the real test for Astra isn’t the benchmark table. It’s whether enterprise customers find that a $50 output token rate translates to fewer human hours and fewer errors on the tasks that actually cost them money.



