TLDRocket
Sign in

Deep Learning Weekly: Issue 463

Deep Learning Weekly Miko Planas Covered by 7 sources

xAI dropped Grok 4.5, a faster, cheaper model, while OpenAI's GPT-Live now lets ChatGPT listen and talk at the same time. Plus a benchmark shows even top LLMs flunk basic data-structure reasoning — turns out talking fast isn't the same as thinking clearly.

Based on reporting by Deep Learning Weekly, Miko Planas — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

xAI's Grok 4.5 is the headline here, and the pitch is efficiency rather than raw brainpower. The company says it runs at 80 tokens per second and needs 4.2 times fewer output tokens than Anthropic's Opus 4.8 to complete SWE-Bench Pro tasks, all while costing just $2 per million input tokens and $6 per million output tokens. That's a meaningful shift in the conversation — for a while, the industry narrative was purely about who could build the smartest model. Now it's about who can build the cheapest smart model, and xAI clearly wants that crown.

OpenAI, meanwhile, is chasing a different kind of efficiency with GPT-Live, a full-duplex voice model that can listen and speak at once instead of taking turns like a walkie-talkie. It hands off harder reasoning or search tasks to GPT-5.5 running quietly in the background, and it's already powering ChatGPT Voice. The bigger story under the hood is what OpenAI's own audit found elsewhere: roughly 30% of tasks in SWE-Bench Pro, a coding benchmark OpenAI had previously championed, turned out to be broken. The company is now walking back its recommendation to use it as a Verified-benchmark replacement — a rare and useful admission that benchmark hygiene matters as much as model hygiene.

There's a broader theme running through this week's research too: the

My take — AI-written commentary, not fact-checked reporting

Grok 4.5's efficiency numbers are the real story, not its intelligence — everyone's racing toward cheap-and-fast because inference cost is now the actual battleground, not benchmark bragging rights. And OpenAI quietly admitting SWE-Bench Pro is 30% broken is the kind of honesty I wish we got more often instead of another triumphant leaderboard screenshot.

Read more about this at: Deep Learning Weekly

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.