Most model releases promise better reasoning. Grok 4.6 makes a more specific bet: that the real bottleneck isn’t raw intelligence but the ability to stay on task across dozens of steps without falling apart. Cursor and SpaceXAI announced the release today, framing it as a direct upgrade to Grok 4.5 with a sharper focus on agentic work, visual projects, and long-horizon coding tasks.
What the benchmarks actually say
Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite score drawn from nine separate benchmarks. That’s a meaningful data point, not just marketing copy. It puts the model in the same tier as OpenAI’s current frontier offering for knowledge work and agentic coding. Whether that translates to real-world performance at the same level is a different question, but the benchmark parity is real and worth taking seriously.
The model is available now inside Cursor and Grok Build, and also through the SpaceXAI API and partners including OpenRouter, Vercel, and Cloudflare. Pricing starts at $2 per million input tokens and $6 per million output tokens. A faster variant runs at twice that price. For the first week, Cursor and Grok Build users get double the included usage.
How it was trained
The training process here is more interesting than usual. Grok 4.6 went through a longer supplemental training run than its predecessor, using curated model-generated data focused on reasoning and technical concepts, plus an updated optimizer and training recipe. That foundation was then used to regenerate supervised fine-tuning trajectories across reasoning, agent tasks, and domains like STEM, software engineering, and general knowledge work. The team used Grok 4.5 itself to generate those trajectories, then filtered out bad traces using model-based checks.
Reinforcement learning was applied across a broad set of agentic tasks, including kernel optimization, web development, and computer-aided design environments. So this isn’t just a model that’s better at answering questions. It’s specifically shaped to operate inside agentic loops.
Where it actually shines
The clearest improvement over Grok 4.5 is in visual and interactive projects. Given a concrete product idea, the model can establish structure and visual layout in a single pass, which matters a lot when you’re iterating quickly. It’s also showing more self-checking behavior on longer tasks, verifying its own work before moving on rather than blindly continuing.
Key improvements over Grok 4.5 include:
- Stronger first-pass quality on visual and interactive applications
- Better sustained performance across multi-step agentic tasks
- Self-testing and verification behavior on longer trajectories
- Broader RL training coverage including kernel optimization and CAD environments
- Widest pre-deployment safety and capability testing suite to date
Why this matters for developers
The competitive context here is important. Anthropic’s Claude 4 Sonnet and OpenAI’s GPT-5.6 are both targeting the same agentic coding space. Google’s Gemini 2.5 Pro is also in this mix. Grok 4.6 entering at benchmark parity with GPT-5.6 Sol, while being available through Cursor natively and priced at $6 per million output tokens, gives it a real shot at being a primary model for developers already inside that ecosystem.
And the 2x usage promo for the first week is a smart move. It lowers the cost of experimentation exactly when curiosity about a new model is highest. For teams evaluating models inside Cursor, this week is a good time to run real tasks against it rather than waiting for the hype to settle.
The agentic coding market is getting crowded fast. But Grok 4.6 isn’t just showing up with a press release. It has a clear training story, specific capability claims, and a distribution strategy that puts it in front of developers where they already work.




