TLDRocket
Sign in

😺 Grok 4.6 is GPT 5.6 level and built for agents that don't quit

The Neuron Eric Gerard Ruiz Covered by 2 sources

Grok 4.6 just launched for long-running AI agents and matches GPT-5.6 on one benchmark. The twist: xAI says it’s built to do the job with fewer turns and lower cost.

Based on reporting by The Neuron, Eric Gerard Ruiz — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

xAI has released Grok 4.6, a new model aimed at long-running agents, coding, research, and other interactive work that doesn’t finish in one neat prompt. Artificial Analysis gave it a score of 61, five points above Grok 4.5 and level with GPT-5.6 Sol overall.

That benchmark number is the shiny part. The practical pitch is stranger and more useful: Grok 4.6 is supposed to stay near the frontier without making the bill explode when an agent needs to think, act, retry, and keep going for hours. That matters because agent systems can look impressive in a single reply and still become expensive the moment the task drags on.

xAI says Grok 4.6 is available now in Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare. Cursor and Grok Build are also offering 2x included usage for the first week. On xAI’s pricing page, Grok 4.6 sits on the $30/month SuperGrok plan, while API pricing starts at $2 per million input tokens and $6 per million output tokens.

Artificial Analysis also put the model at $0.84 per task and said it landed on its cost-performance frontier across every agentic evaluation in the index. One comparison stood out: on AA-Briefcase, a Grok reached Fable 5-tier while averaging about 53 turns and 0.5B input tokens, versus roughly 103 turns and 2.0B for Claude Opus 5 Max. Less back-and-forth means less token burn, and in agent land that’s often the whole game.

Musk says Grok 4.7 is already coming in 3 to 4 weeks and is supposedly much better than 4.6, with a large dose of SpaceX company data behind it. So this version may not get to enjoy a very long victory lap.

My take — AI-written commentary, not fact-checked reporting

The market keeps acting like agent quality is the headline, when the ugly little invoice is usually the real product. A model that gets the job done in fewer turns is the kind of boring advantage that actually survives contact with teams, budgets, and Monday morning. The hype cycle loves superlatives; operations just wants fewer retries.

Read more about this at: The Neuron

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.