TLDRocket
Sign in

SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work

MarkTechPost Michal Sutter Covered by 3 sources

SpaceXAI shipped Grok 4.6, a post-training upgrade with 500K context. It’s built for long agent runs and coding, but there’s still no open-weights path.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

SpaceXAI has released Grok 4.6, and the interesting part is what didn’t change: the base model. This is a post-training upgrade to Grok 4.5, not a bigger foundation model. The company spent the lift on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning in agentic environments that reward models for staying on task across many steps without drifting.

The result is a model with a 500,000-token context window, text and image input, and text-only output. It also adds a new xhigh reasoning-effort setting above the low, medium, and high options Grok 4.5 had. SpaceXAI says Grok 4.6 is live today in Cursor and Grok Build, generally available through the xAI API as grok-4.6, and routable through OpenRouter, Vercel, and Cloudflare. There’s no open-weights release and no self-hosting path, which means air-gapped deployments are off the table.

On the company’s launch table, Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max. It also posts gains on business-style tasks: GDPval-AA v2 at 1753 Elo, AA-Briefcase at 1577, and Harvey LAB. But the coding numbers are less flattering. DeepSWE v1.1 comes in at 65.9%, Terminal-Bench v3.0 at 26%, CursorBench v3.2 at 69.9%, FrontierCode v1.1 Extended at 61.3%, and APEX-Agents at 57.5%.

That weaker showing matters because SpaceXAI is pitching this as a model for repository-wide refactors, migration agents, research over huge corpora, scaffolding from product briefs, GPU kernel optimization, and other knowledge-heavy work. The company also says longer trajectories led the model to test and verify its own work more often. That’s a vendor observation from internal testing, not an independent measurement, and the launch table itself is the cleaner signal.

Pricing stays in the usual xAI neighborhood: $2 / $0.50 / $6 per million tokens below 200K prompt tokens, then $4 / $1 / $12 above that. A faster variant is mentioned at double the price, but no separate model ID is published. Cursor and Grok Build are also offering 2x included usage for the first week. For teams that want to try it, SpaceXAI says to set a prompt_cache_key, or the x-grok-conv-id header in Chat Completions, or pay full input price when cache hits get messy.

My take — AI-written commentary, not fact-checked reporting

This is the familiar frontier-model playbook now: turn the crank on post-training, call it a new era, and hope the benchmarks cooperate. The catch is that the model looks strongest exactly where enterprise buyers like to talk and weaker where engineers actually sweat, which is a very modern kind of disappointment. No open weights, no self-hosting, and a procurement headache baked right in — efficient, if the goal is to sell access rather than trust.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.