Grok 4.7 Pushes SpaceXAI Into the Top 4 of AI Labs
Trending Topics Jakob Steinschaden ● Covered by 6 sources
SpaceXAI’s Grok 4.7 is out, and it lifts the lab into the top 4 on a major AI benchmark. Big jump in coding and long-work tasks, but it also burns a lot more tokens to get there.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
SpaceXAI has shipped Grok 4.7, and the release gives the company its best showing yet in the Artificial Analysis Intelligence Index. On the index, the model scores 46 points, up two from Grok 4.6, which is enough to put SpaceXAI fourth among frontier-model labs when the test is run at the highest reasoning setting, xhigh.
The company is leaning hard on two strengths: agentic knowledge work and coding. It says Grok 4.7 is built on a larger base model than Grok 4.6, trained with a longer reinforcement learning run, and pushed on harder tasks that can take hours to finish. That seems to have paid off in self-checking and long-context handling. It was also trained to understand Grok Bot natively, which should help in conversational work and general knowledge tasks.
The clearest progress shows up in long-horizon benchmarks. On AA-Briefcase, Grok 4.7 scores 1,657 Elo, up from 1,546 for Grok 4.6 in the high configuration. Only Claude Opus 5 and Claude Fable 5.1 sit above it, with Fable 5.1 just ahead at 1,678. GDPval tells a similar story: 1,695 Elo for Grok 4.7 versus 1,605 for Grok 4.6. The model also improves on coding, where Grok 4.7 plus Grok Build scores 56 points on the Artificial Analysis Coding Agent Index, enough to move ahead of GPT-5.6 Sol.
But the gains are uneven. SpaceXAI’s own domain tests put Grok 4.7 ahead in electrical engineering and legal work, yet behind in clinical reasoning. On HealthBench Professional it scores 56.7 percent, trailing GPT-5.6 Sol at 60.5 percent and Fable 5.1 at 62.1 percent. And there’s a cost to the progress: Grok 4.7 uses about 81,000 output tokens per Intelligence Index task, far above Grok 4.6’s 36,000. The company says the model is available now through Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers and cloud platforms, with pricing unchanged and a faster variant available at twice the price.
My take — AI-written commentary, not fact-checked reporting
This is the familiar AI trade: better results, heavier compute bill, and a lot of marketing about safety stapled on top. SpaceXAI can call Grok 4.7 its strongest model yet, but the real story is that the frontier keeps rewarding models that are expensive to run and hard to ignore. Very efficient, if your hobby is paying for tokens.
Read more about this at: Trending Topics