xAI is not playing it quiet with Grok 4.7. The company announced the model as its most capable release yet for coding and knowledge work, and the pricing is the part that will get developers to stop scrolling. Starting at $2 per million input tokens and $6 per million output tokens, it matches Grok 4.6’s pricing while delivering meaningfully better performance. For context, that puts it well below what you’d pay for comparable output from GPT-4o or Claude Sonnet in equivalent workloads.
What actually changed under the hood
Grok 4.7 is built on a new, larger base model than its predecessor. That alone is worth noting, but the more interesting detail is the training approach. xAI ran a longer reinforcement learning process on a harder mix of tasks, specifically weighted toward problems that take many hours to complete. The practical result is a model that is better at sticking with difficult tasks, checking its own outputs, and managing longer context windows without degrading.
The model was also trained to natively understand the Grok Bot harness. That matters for conversational workflows and general knowledge tasks where context continuity is important. It’s not just a coding upgrade, it’s a more coherent model across use cases.
On CursorBench 4.0, which stresses long-running coding tasks specifically, Grok 4.7 sits at the frontier on price-performance. That benchmark is increasingly relevant as AI coding tools shift from autocomplete toward multi-step, multi-file tasks. So this isn’t a benchmark cherry-pick. It reflects real developer workflows.
Beyond code: documents, presentations, and professional tasks
xAI also tested Grok 4.7 on GDPval and AA Briefcase, two benchmarks where AI is asked to handle tasks typically done by lawyers, nurses, and financial analysts. Grok 4.7 improves on Grok 4.6 across both and performs comparably to other frontier models. This positions it less as a pure coding tool and more as a general-purpose model for knowledge workers, which broadens its appeal beyond just dev teams.
A rebuilt safety stack that doesn’t over-restrict
The safety story here is actually more nuanced than most model releases. Grok 4.7 ships with an entirely new safeguard stack, and xAI claims it’s the strongest model they’ve tested on refusals and jailbreak resistance. On LatchBio’s biosafety benchmark, it scores 62.4%, which leads the field.
But what makes this interesting is the dual-use calibration. On HackerBench v0.3, Grok 4.7 blocks only 3.3% of risky dual-use prompts while keeping refusal rates low for legitimate security work. That’s a hard balance to strike, and most models either refuse too aggressively or not enough. xAI is also giving select cybersecurity partners invite-only access to red-team capabilities for defense research, which signals they’re taking the responsible deployment angle seriously rather than just citing benchmark numbers.
Availability and pricing
Grok 4.7 is available now across several surfaces:
- Cursor and Grok Build
- The Grok API directly
- Third-party coding harnesses
- Model routers and cloud platforms
Standard pricing is $2 per million input tokens and $6 per million output tokens. A fast variant is also available at twice the output speed for twice the price. For teams running high-volume coding pipelines, the base tier is competitive enough to justify a real evaluation against Claude and GPT-4o.
The broader trend here is that the mid-tier of frontier models is getting genuinely crowded. Mistral, Anthropic, and Google are all pushing capable models in this price range. Grok 4.7 enters that competition with credible benchmark performance and a pricing structure that at minimum earns it a spot in the evaluation shortlist.



