A 70.6% score on Terminal-Bench 4.0 versus 10.3% for its predecessor is not a incremental improvement. That’s a different class of model. Anthropic has announced Claude Sonnet 5.5, the second model in the Claude 5.5 family, and the numbers suggest it punches significantly above the Sonnet tier in agentic coding tasks.
Sonnet 5.5 is positioned between Opus 5.5 (for complex, judgment-heavy work) and the upcoming Haiku 5.5 (for high-volume, cost-sensitive applications). Its target use case is well-defined tasks: fixing bugs, drafting documents, building slides and spreadsheets, and iterative coding work where speed matters. Pricing holds at $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. But Anthropic says the model typically needs fewer tokens to complete the same work, which puts the real-world cost reduction at up to 30% per task.
What the benchmarks actually show
The jump in agentic coding is the headline. On CursorBench 4.0, which tests models on real Cursor coding sessions, Sonnet 5.5 scores 55.5% compared to Opus 5.5’s 57.8%. On GDPval-AA, a knowledge work benchmark spanning real-world occupational tasks, Sonnet 5.5 scores 1844 versus Opus 5.5’s 1846. That’s effectively the same result at a lower price point, depending on effort level.
Speed is also a real differentiator here. Sonnet 5.5 generates outputs more than 30% faster than Sonnet 5, which matters in iterative workflows where you’re running the model repeatedly. Early testers at companies like CodeRabbit noted that the model batches tool calls more efficiently than Sonnet 5, reducing steps and cutting token usage in the process.
Coding performance is where it stands out
On FrontierCode 1.1 at High effort, Sonnet 5.5 scores 10 points higher than Sonnet 5 at the same setting, at roughly one-fifteenth of the cost per task. Epic Games reported that the model handled tens of thousands of lines of code for gameplay system architecture in early testing, with multi-hour tasks completed without heavy prompt engineering. Base44 found that across 118 real app builds, Sonnet 5.5 reached the same quality bar as Opus 5 in an average of 3.6 iterations per build, compared to 7.7 for Opus 5.
So this is not just a benchmark story. The efficiency gains appear to hold in production environments, which is where most developers will actually care.
Key capabilities and specs
- Terminal-Bench 4.0: 70.6% (vs. 10.3% for Sonnet 5, 66.4% for Opus 5.5)
- OSWorld 2.1 computer use: 80.1% partial (vs. 57.0% for Sonnet 5)
- Humanity’s Last Exam with tools: 64.5% (vs. 54.9% for Sonnet 5)
- Output speed: 30%+ faster than Sonnet 5
- Pricing: $2/million input tokens, $10/million output tokens, $0.20/million for cache reads
- First Sonnet model to include Opus-level cyber safeguards and fallbacks
Safety and alignment additions
Because Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, Anthropic made a specific decision to include the cyber safeguards and fallbacks previously reserved for their most capable models. This is the first Sonnet-tier model to get that treatment. Biology safeguards match Sonnet 5’s. According to Anthropic, both sets of restrictions target a narrow set of high-risk requests and don’t affect routine software development or standard life sciences work.
Why this matters for developers choosing between Claude models
The honest read here is that for most production workloads, Sonnet 5.5 just made the choice easier. If you were using Sonnet 5, there’s no reason not to upgrade. If you were paying for Opus 5.5 on tasks that don’t require deep open-ended reasoning, Sonnet 5.5 at medium or high effort may be sufficient at a lower cost per run. Opus 5.5 still leads on complex, open-ended work requiring sustained judgment, and Anthropic is clear about that distinction.
But for coding agents, document workflows, and fast iterative tasks, Sonnet 5.5 closes enough of the gap that the price difference becomes hard to justify on those use cases. Claude Haiku 5.5 is still weeks away, but when it arrives, Anthropic will have a complete three-tier lineup in the 5.5 family. That’s a competitive stack against OpenAI’s GPT-4o/o3 structure and Google’s Gemini tiering.



