Deep Learning Weekly: Issue 475
Deep Learning Weekly Miko Planas ● Covered by 18 sources
OpenAI, Anthropic and Nvidia all shipped new AI gear this week. The big theme: cheaper, faster agents with more guardrails and memory.
Based on reporting by Deep Learning Weekly, Miko Planas — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
This week’s issue is basically a snapshot of where the agent race has gone: speed, price, and control. OpenAI pushed out GPT-6.1 Sol, which it says runs at Astra-like performance for one-fifth the price, and added an Ultrafast tier that reaches 300 tokens per second. It also says the model beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 on AutomationBench. That’s not just benchmark garnish. It’s a reminder that the next battleground is how much work an agent can do before the bill starts to hurt.
Anthropic answered with Claude Sonnet 5.5, which it says cuts per-task costs by 30% thanks to faster speeds and fewer tool calls. The model scores 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1, and comes close to Opus 5.5 on GDPval-AA. OpenAI, meanwhile, is leaning into the always-on agent story with Dots and ChatGPT Space. Each Dot gets its own cloud computer and browser, runs on GPT-6 Astra, and is fenced in by Custom Rules, with passwords left to humans. Helpful, if you like coworkers who cannot reach the secrets.
The infrastructure side is getting more serious too. Nvidia’s Open Agent Safety Platform moves access control and monitoring onto BlueField-4 DPUs, while Cohere’s Embed 5 splits indexing and querying across Pro and Fast without forcing a re-index. Black Forest Labs’ FLUX 3 Action brings an open-weight 7B world-action model for robot control, and AMD says it will buy World Labs for $8.2 billion, naming Fei-Fei Li as executive vice president and chief scientist.
The research section keeps circling the same pain points: memory, search, and evaluation. JitMem argues for deferring memory curation until read time, and says that approach beat the strongest baseline by 16.2, 16.3, and 3.9 success-rate points across ALFWorld, WebShop, and τ2-bench. IterSynth splits planning from synthesis for deep search and posts a 50.7 average on five benchmarks, while Jev-as-judge beats GPT-4o-mini on cost and speed in a 1,000-turn comparison. The message across the whole issue is pretty blunt: agents are getting cheaper and more capable, but only the ones wrapped in better memory and tighter controls are looking ready for real use.
My take — AI-written commentary, not fact-checked reporting
The industry still loves pretending bigger models are the story, but this issue says the grown-up work is everything around them: memory, evals, controls, and cost. OpenAI and Anthropic can keep one-upping each other on speed, but the boring plumbing will decide who actually gets used. Agents without guardrails are just expensive interns with browser tabs.
Read more about this at: Deep Learning Weekly