TLDRocket
Sign in

Deep Learning Weekly: Issue 474

Deep Learning Weekly Miko Planas

Anthropic cut Opus 5.5’s price and OpenAI fired back with GPT-6 Sol and Luna. The week’s AI race is now about cost, speed, and fewer mistakes.

Based on reporting by Deep Learning Weekly, Miko Planas — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

This week’s Deep Learning Weekly reads like a pricing war with benchmarks attached. Anthropic shipped Claude Opus 5.5 at $4/$20 per million tokens, 40% below Opus 5, while saying it gets 66.4% on Terminal-Bench 4.0 and produces output more than 30% faster. OpenAI answered within hours with GPT-6 Sol and Luna, pricing the series at half the 5.6 tier and claiming Sol makes about half as many factual mistakes as its predecessor.

The open-weights side had its own flex. Xiaomi launched MiMo-V2.6-Pro under an MIT licence, calling it the top open weights model in the world. It’s a 1.02-trillion-parameter MoE with 42 billion active parameters, ties Grok 4.7 at 46 on Artificial Analysis, and costs $0.435 per million input tokens.

Cost, though, keeps sneaking back into the story. xAI’s Grok 4.7 raises Terminal-Bench 4.0 from 20.3% to 38.0%, but it also uses 125% more output tokens than 4.6, which pushes the cost to $3.74 per Intelligence Index task. PrismML is trying the opposite trick with Bonsai 2 27B, a ternary Apache-2.0 model compressed to 5.9GB that still keeps roughly 98.2% of capability and runs at 143 tokens per second on a consumer GeForce 5090.

There’s also a clear shift in where the work is happening. Anthropic opened a Bay Area wet lab where Claude-driven robots run experiments through a Model Hardware Standard protocol with the Howard Hughes Medical Institute. AWS launched Strands Harness, an open-source agent harness it says is 26% more efficient and 77% cheaper than Claude Code on identical tasks. And in the papers, RRSI argues that harnesses, not just backbone models, drive capability — while an empirical study of coding harnesses and an Agensh system scaling to 1,024 agents both point to the same unglamorous truth: orchestration is the product now.

The most useful part of this issue is how little faith it places in vibes. Prices are down, token use is up, and the best gains keep coming from better harnesses, better context handling, and more careful evaluation. The model race is still loud, but the bill is written somewhere else.

My take — AI-written commentary, not fact-checked reporting

The smart money is clearly moving from “which model wins?” to “which harness wastes the least money doing it?” That’s the part the glossy demos always skip, because competence is boring and billing is rude. Open models, open harnesses, and reproducible evals are the only things that make this market less of a magic trick.

Read more about this at: Deep Learning Weekly

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.