TLDRocket
Sign in

SpaceXAI releases Grok 4.7 for coding and knowledge work

TestingCatalog AI News Covered by 9 sources

SpaceXAI launched Grok 4.7, a new model for coding and knowledge work. It’s cheaper than many rivals, works longer on hard jobs, and comes with tighter safety limits.

Based on reporting by TestingCatalog AI News — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

SpaceXAI has released Grok 4.7, calling it its most capable model yet for coding and knowledge work. Access is already live through Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms. Grok Build also lets people try it for free.

The company is pitching the model as a faster, lower-cost alternative to other frontier systems. Pricing starts at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. There is also a fast variant that doubles output speed and doubles the price.

Under the hood, Grok 4.7 uses a larger base model trained with a longer reinforcement learning cycle and a harder mix of tasks, including work that can stretch for many hours. SpaceXAI says that makes it better at staying on a difficult assignment, checking its own output more carefully, and handling long context. It was also trained natively on the Grok Bot harness, with an eye toward chat and general knowledge work.

The benchmarks show the gains are real, though not universal. Grok 4.7 reached 46.3% on CursorBench 4.0, up from 40.4% on Grok 4.6, and hit 71.0% on DeepSWE v1.1 at high effort. It also posted 64.0% on EEBench, 1,657 on AA Briefcase v1.1, 38.0% on Terminal-Bench 4.0, and 19.6% on the Harvey Legal Agent Benchmark. But its 56.7% HealthBench Professional score trailed GPT-5.6 Sol Max and Fable 5.1 Max, which is a reminder that one model does not own every category.

Safety is a bigger part of this release than the usual launch chatter suggests. SpaceXAI says Grok 4.7 uses a completely new safeguard stack and is its strongest model yet for refusals and jailbreak resistance. It scored 62.4% on LatchBio’s biosafety benchmark and let only 3.3% of risky dual-use prompts through on HackerBench v0.3, while still rarely blocking legitimate security work. A few cybersecurity partners are also getting invite-only access to its red-team features for defense research.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of release: more useful, less theatrical. The interesting bit isn’t the benchmark victory lap, it’s the combo of long-task endurance, self-checking, and tighter refusals — the stuff that actually matters when people pay for models to do work, not perform for a demo. The AI market could use fewer fireworks and more models that don’t fall apart halfway through a job.

Read more about this at: TestingCatalog AI News

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.