TLDRocket
Sign in

[AINews] not much happened today

Latent Space Covered by 4 sources

OpenAI's coding agents got slammed with demand this week, and a startup squeezed a 27B model down to 3.9GB. The real story: the harness around a model now matters as much as the model itself.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sam Altman spent the week bragging and sweating in equal measure. Codex plus ChatGPT Work usage grew 2.5x in a single week, he said, and then followed up that demand for GPT-5.6 Sol is "insane" enough that infrastructure might not keep pace. JetBrains responded by making Codex its recommended agent. OpenAI ran multiple usage resets to keep the lights on. This is the kind of growth curve that looks great in a slide deck and terrifying in an on-call rotation.

But the more interesting thread running through the week wasn't raw usage, it was what happens once you actually try to run these agents for hours at a time. Swyx flagged that stale agents.md instruction files can act like self-inflicted prompt injection, quietly derailing long-running tasks for hours before anyone notices. LangChain responded by extending its tracing tools from Codex to Cursor, Copilot, Pi, and OpenCode, exposing tool calls and token usage that used to be invisible. Teknium shipped Hermes updates that let it parallelize tool calls. Andy Konwinski put the underlying thesis plainly: companies that turn their own workflows into evals and environments may end up with a sturdier edge than the ones just throwing more compute at bigger models.

Meanwhile the compression side of open models kept getting more aggressive, and more useful. PrismML released Bonsai 27B, built on Qwen 3.6 27B, in a ternary version at 5.9GB and a 1-bit version at just 3.9GB, both under Apache 2.0. That's a 27-billion-parameter model small enough to run on a phone or a single consumer GPU, and a demo showed it running agentic, tool-using tasks on an RTX 5090 without falling apart. Tencent's Hunyuan team pushed in a similar direction with Hy3, a 295B-parameter model offered in 1-bit and 4-bit forms that can be served on one GPU through llama.cpp. Local inference stopped being a novelty a while ago; this week it looked closer to a genuinely viable path for serious agentic work.

Elsewhere, Perplexity open-sourced WANDR, a 500-task benchmark for wide-and-deep research built from real production tasks, backed by over 170,000 source-linked records, and designed to re-check citations against live web pages rather than a frozen answer key. Sakana AI, working in a completely different register, published research in Nature Communications on "Smart Cellular Bricks": identical cubes that talk only to their neighbors yet can detect missing pieces with 95% accuracy across six directions and regrow damaged structures, a method that scaled to more than 18,000 cubes in simulation. And in a smaller, stranger corner of the field, someone posted a micro-drone autonomously intercepting a flying moth, pitched as an early step toward mosquito control.

Calling this "not much happened" is technically accurate and also very funny, because between demand spikes threatening OpenAI's infrastructure, models shrinking to phone-friendly sizes without losing their agentic chops, and cube-shaped robots self-organizing without a central brain, it was a fairly loaded week for a slow news day.

My take — AI-written commentary, not fact-checked reporting

The interesting fight right now isn't which lab has the biggest model, it's who builds the tooling that keeps a model from wandering off task for hours because of a stale instructions file. OpenAI's usage numbers are real, but so is the infra strain that comes with them, and that tension is worth watching more than any benchmark score. Meanwhile the quiet compression race, squeezing a 27-billion-parameter model onto a phone, deserves more attention than it's getting; that's the kind of shift that actually changes who gets to run these systems, not just how impressive the demo looks.”}

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.