Grok 4.6 lands in Cursor with 500K context and a $6 output price
Curated by the Inblix editorial team
SpaceXAI shipped Grok 4.6 today, and the interesting story isn’t a bigger model — it’s what happens when you keep the foundation frozen and spend your budget on post-training instead. The company ran a longer supplemental training run than Grok 4.5 received, regenerated supervised fine-tuning trajectories using Grok 4.5 itself, then pushed the model through reinforcement learning in agentic environments covering knowledge work, coding, web development, CAD, and kernel optimization. Same base, better behavior.
The headline number: 61 on the Artificial Analysis Intelligence Index, up five points and tied with GPT-5.6 Sol Max. That’s real progress. But look at the losses before the wins. DeepSWE v1.1 lands at 65.9% — an 11.9-point generational jump, yet still behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 hits 26%, nearly double Grok 4.5’s 15.7%, and still sits last among the four listed models. Meanwhile, the bolded victories on GDPval-AA v2 and AA-Briefcase fall inside published confidence intervals. Those are statistical ties, not leads. The comparison set also conveniently omits Claude Opus 5, which currently tops that index.
The deployment story is pragmatic. Grok 4.6 is live in Cursor on all plans and is the default in Grok Build, which means seed-stage teams and indie developers can adopt it without harness work. Mid-market engineering orgs get API-only integration with mTLS and batch processing documented. Regulated enterprises should pilot first — SpaceXAI’s brand history is reportedly a live procurement question in multiple buying committees. No open weights, no self-hosting, no air-gapped path.
Pricing runs $2 per million input tokens under 200K prompt tokens, doubling to $4 above that threshold, with output at $6 to $12 depending on context length. A faster variant exists at double the price but no separate model ID has been published. One operational detail worth flagging: teams need to set prompt_cache_key or the x-grok-conv-id header. Skip it and requests scatter across servers, cache hits collapse, and you pay full input price. The new xhigh reasoning-effort level is the other lever — SpaceXAI reports more self-testing and verification on longer trajectories, though that’s a vendor observation from internal testing, not an independently measured result.
💡 Key Takeaways
- Grok 4.6's gains come entirely from post-training — a longer supplemental run, regenerated SFT trajectories, and RL in agentic environments — not from a larger base model.
- The model ties GPT-5.6 Sol Max at 61 on the AA Intelligence Index but trails on the coding benchmarks engineering teams actually care about, including DeepSWE and Terminal-Bench.
- A new xhigh reasoning-effort level is available, but the claim of increased self-testing on long trajectories is vendor-reported, not independently verified.
- Without setting prompt_cache_key, teams will scatter requests across servers and lose cache hits, effectively paying full input price on every call.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.