Grok 4.7 was built to work for hours. It still fails most of the time.
The New Stack Amanda Caswell
SpaceXAI’s Grok 4.7 was trained for long coding jobs, and it got better at them. It still trails the best rival on a key test, so the hard part isn’t fixed.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
SpaceXAI says Grok 4.7 was built for a very specific kind of pain: coding work that runs for hours, picks up lots of context, and can go off the rails if the model misses one bad decision. The new version, released on Sunday, was trained with a longer reinforcement learning run that leaned into harder tasks, including jobs that take many hours. SpaceXAI says that approach also improved self-verification and long-context handling.
The numbers show a real jump. Grok 4.7 scored 38.0% on Terminal-Bench 4.0, up from 20.3% in Grok 4.6. On CursorBench 4.0, which focuses on longer coding workflows inside an editor, it rose from 40.4% to 46.3%. And on AA Briefcase v1.1, a test of multi-hour professional work, it moved from 1,546 to 1,657.
But the gap at the top is still there. On Terminal-Bench 4.0, Anthropic’s Claude Fable 5.1 scores 57.9% on the independent leaderboard, well ahead of Grok 4.7. That matters because long-horizon agent work is exactly where mistakes compound. A bad assumption early on can poison everything that follows.
SpaceXAI is also pushing harder on the plumbing around the model. Grok 4.7 was trained to understand the Grok Bot harness natively, which means the model is being taught alongside the tools, terminal formatting, execution feedback, and next-step logic it will use. OpenAI took a similar route last week by turning the Codex harness into the Agents API. The direction is clear: the model is no longer the only thing being tuned.
SpaceXAI hasn’t said what changed under the hood to improve context management or self-verification, and it hasn’t broken out how much of the gain comes from harness-specific training. Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens, which may make long runs easier to stomach. Reliability still decides whether those hours buy progress or just expensive wandering.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI that matters now: not flashy demos, but models that can sit there, remember what they did, and not trip over their own shoelaces. The industry keeps pretending better tools alone will save agents, but the real contest is between better models and better systems around them. Right now, the scoreboards still say the job is unfinished.
Read more about this at: The New Stack
Related stories
SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work
MarkTechPost · 1 month ago ·
16