TLDRocket
Sign in

[AINews] AMD buys Taalas

Latent Space

AMD is buying Taalas, the startup betting that etching whole LLMs directly into custom chips beats running them on GPUs. Meanwhile Meta's small Muse Spark 1.2 model is punching way above its price, and OpenAI just merged its two chat modes into one.

Lisa Su just put real money behind a bet that a lot of people in the industry have been quietly making jokes about for the past year: that the future of inference isn't GPUs running ever-larger models, but ASICs with the model's weights baked directly into silicon. AMD's acquisition of Taalas is the clearest signal yet that at least one major chipmaker thinks etched LLMs are more than a science project. Latent Space flagged Taalas months ago as one to watch, and skeptics on the Baseten podcast pushed back hard on the idea. Su, evidently, wasn't swayed by the skepticism.

While that deal was landing, the rest of the AI world kept moving at its usual dizzying pace. Meta's Muse Spark 1.2 jumped from off-the-radar to top-five on the Vals Index, running at $0.69 per test — roughly a third the cost of Kimi and a tenth the cost of Opus 5 or Fable. It also became the first model to crack 60% on Finance Agent v2, undercutting Opus 5's price by nearly 7x while running twice as fast. Meta says Muse Spark-family models hit gold-medal territory across five STEM Olympiads without using any tools, which immediately reignited the argument over whether that's genuine reasoning or clever multi-agent orchestration doing the heavy lifting.

OpenAI, for its part, simplified its own product line by killing the split between Instant and Thinking modes. GPT-5.6 Sol now handles both jobs in one chat surface for paid users, with a slider to dial reasoning effort up or down, and OpenAI claims a 68% drop in factual errors on a high-stakes eval spanning finance, medicine, and law. Free users aren't left out either — unlimited GPT-5.6 Luna chats are rolling out, and ARC Prize confirmed that an 80% price cut on Luna came with zero capability loss. OpenAI also shipped Agent Plugins, a shared standard for packaging skills and MCP configs across Codex, Cursor, GitHub Copilot, and more.

Underneath all of this, the infrastructure layer is quietly turning into the real battleground. Cloudflare's Kitesurf strips browser automation down to a stateless, Chromium-free process built for agents rather than humans. Weaviate baked an MCP endpoint straight into its REST API. And François Chollet's argument that harness-heavy inference systems are secretly neurosymbolic — not pure neural end-to-end programs — is turning into a real engineering debate rather than an academic one, because routing and orchestration choices are visibly changing what these systems can do.

Google DeepMind, meanwhile, published WeatherNext 2 in Nature, open-sourcing weights that reportedly buy an extra day of hurricane forecast lead time — described internally as a decade's worth of progress in one release. It's a reminder that amid all the chatbot benchmark wars, some of this technology is still just quietly getting better at predicting where a storm will make landfall.

My take

AMD buying an etched-LLM startup is a bigger deal than the modest headline suggests — it's a hardware giant hedging that GPUs won't stay the default inference substrate forever, and that should worry Nvidia more than any single model release this week. Everyone's obsessing over which chatbot topped which leaderboard, but the actual structural shift is happening one layer down, in silicon and routing decisions nobody outside infra teams is paying attention to. Watch the chips, not the chat windows.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.