TLDRocket
Sign in

SpaceXAI trained Grok 4.6 on something most AI labs throw away

The New Stack Amanda Caswell Covered by 3 sources

SpaceXAI shipped Grok 4.6 just weeks after 4.5. It was trained to learn from bad code, and that may matter more than flashy first tries.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

SpaceXAI pushed out Grok 4.6 on Wednesday, less than a month after Grok 4.5. The pitch is not just that it writes code. The company says it can research unfamiliar topics, move through large codebases, and take a product idea all the way to a working app.

The bigger shift is in how it was trained. SpaceXAI says Grok 4.6 got a longer supplemental training run than 4.5, mixing model-generated reasoning, technical material and engineering data. It also changed the optimizer and the training recipe used to update the model’s weights. Then the company used Grok 4.5 to rebuild supervised fine-tuning trajectories across different reasoning settings, agent harnesses and domains such as STEM, software engineering and knowledge work.

Some of those trajectories were thrown out after model-based checks flagged them as problematic. Reinforcement learning then extended the process into general coding, kernel optimization, web development and computer-aided design. The goal was to reward the model for finishing the larger task, not just for producing code that looks good at a glance. SpaceXAI says that made Grok 4.6 more likely to pause on long jobs, check its work, and correct itself before moving on. It also says the model produced stronger early versions of visual and interactive apps.

The benchmark picture is more mixed than the headline numbers suggest. Grok 4.6 beat 4.5 on the tests the company highlighted, including CursorBench v3.2, DeepSWE v1.1, FrontierCode v1.1 Extended, Terminal-Bench v3.0, APEX-Agents and APEX-SWE. But it still trailed Anthropic’s Fable 5 Max on several of those tests, and it also fell behind on Terminal-Bench by a wide margin. On Artificial Analysis Intelligence Index, it rose five points to 61, tying GPT-5.6 Sol Max but still behind Claude Opus 5 and Fable 5.

Pricing is part of the story too. Grok 4.6 keeps the same API rates as 4.5: $2 per million input tokens and $6 per million output tokens. That makes the standard model less than half the API cost of GPT-5.6 Sol Max before you even count retries, tool calls or all the other ways agents can burn money. SpaceXAI is betting that the real win is not the cheapest token, but the job that ends sooner.

My take — AI-written commentary, not fact-checked reporting

This is the right fight to be having. Training models to notice their own mistakes is a lot more useful than worshipping the first answer they cough up. The industry’s habit of bragging about token prices while ignoring task costs has been a neat little accounting trick, and it’s overdue for retirement.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.