A 680,000-line code migration completed in under a day. That’s not a marketing claim buried in fine print — it’s what an early tester actually did with Claude Opus 5.5. Anthropic has announced its new flagship model, and the combination of higher capability and lower cost is the most interesting thing happening in the frontier model space right now.
What Opus 5.5 actually is
Opus 5.5 is the first release in Anthropic’s Claude 5.5 family, sitting above Opus 5 in capability while costing 40% less to run on typical workloads. That efficiency gap matters a lot for anyone running agentic pipelines or long coding sessions, where token costs pile up fast. Cache reads, which make up the bulk of costs in agentic and coding work, drop to $0.20 per million tokens — down 60% from Opus 5. Input tokens are $4 per million and output tokens are $20 per million, both down 20%.
Speed is also up. Opus 5.5 generates output more than 30% faster than Opus 5. And for users who need even more throughput, a fast mode is available in Claude Code and the Claude Platform at up to 2.5x speed, priced at $8 per million input tokens and $40 per million output tokens.
Benchmark results and where they matter
On the benchmarks Anthropic published, Opus 5.5 leads across agentic coding, computer use, and knowledge work. On Terminal-Bench 4.0, it scores 66.4% against GPT-6 Astra’s 57.9% and Opus 5’s 52.3%. On FrontierCode v1.1, it hits 54.4%, ahead of Astra at 53.3%. On OSWorld 2.0 for computer use, it scores 81.8% partial, just above Fable 5.1 at 80.7%.
But Anthropic is honest about benchmark limitations at this level. The internal assessment is that the real-world gap between Opus 5.5 and Claude Fable 5.1 is narrower than raw scores suggest. Where the advantage is clearer is efficiency — fewer tokens per task, lower price per token, net result of 40% lower costs. On FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal-Bench 4.0, it matches Astra for about 40% of the cost.
Coding performance in practice
The real-world coding results are striking. One early tester audited and fixed a 200,000-line codebase in under three hours. Opus 5 took over 20 hours on the same job and used 2.5x as many tokens. In an internal test, Anthropic asked both Opus 5.5 and Fable 5.1 to translate HAProxy from C into Rust. Both passed nearly all of HAProxy’s regression tests. Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
So for code-heavy workflows, this isn’t a marginal upgrade. The cost and speed differences at scale are significant enough to change how teams structure their usage.
Safety and alignment testing
Anthropic is positioning safety as a genuine feature here, not an afterthought. Opus 5.5 achieves the highest scores yet on Anthropic’s automated behavioral audit, which tests the model across thousands of simulated scenarios. It’s less likely to take hard-to-reverse actions, better at staying within defined boundaries, and more resistant to prompt injection than Opus 5.
Because Opus 5.5 is comparable to Claude Fable 5.1 in biology and cybersecurity capability, it ships with equivalent safeguards. Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use it for biology research. A Cyber Verification Program for security practitioners is coming in the next few weeks.
- Automated behavioral audit scores: highest of any Anthropic model to date
- Reduced likelihood of taking irreversible actions without authorization
- Stronger resistance to prompt injection vs. Opus 5
- Broader alignment testing covering longer tasks, impossible tasks, and real-incident scenarios
- External evaluation by Frontier Design and METR before release
Availability and what’s next
Opus 5.5 is available now. Anthropic is also raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans, and giving subscription users a rate limit reset they can save and use on demand. Claude Sonnet 5.5 and Claude Haiku 5.5 are coming in the next few weeks with similar improvements across performance, efficiency, and safety.
For developers and teams running serious workloads, the cost reduction alone is worth evaluating. But the combination of lower costs, faster output, stronger safety scores, and genuine benchmark leads over GPT-6 Astra on most coding tasks makes Opus 5.5 the most competitive Anthropic has looked against OpenAI’s frontier models to date.




