The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
Substack Jesus Rodriguez ● Covered by 7 sources
OpenAI, Meta, and xAI all dropped new AI models and tools last week — GPT-5.6, Grok 4.5, Muse Spark 1.1, plus agent apps. The chatbot is turning into a full work engine, not just a chat window.
Based on reporting by Substack, Jesus Rodriguez — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Last week felt less like a model release cycle and more like an infrastructure land grab. OpenAI shipped GPT-5.6 split into three variants — Sol, Terra, and Luna — each tuned for a different intelligence-per-dollar tradeoff, alongside a feature that lets the model write its own coordination programs to juggle tools and subagents. That's not autocomplete anymore. That's something closer to a scheduler.
The interface layer moved just as fast. GPT-Live ditches the old turn-based voice assistant model for full-duplex audio that can listen and talk at once, interrupt, or go quiet while reasoning runs in the background. ChatGPT Work, meanwhile, is built to sit inside a project for hours, touching files and apps and spitting out finished slide decks or spreadsheets instead of just answers. The product being sold has quietly changed from a reply to a completed artifact.
Meta and xAI answered in kind. Muse Spark 1.1 pairs a million-token context window with computer-use skills and a genuinely clever trick: it compacts long sessions without losing the state it'll need later, and it decides on the fly whether to script an action or just click through an interface itself. Meta also launched a paid Model API, a clear signal it wants to sell metered intelligence, not just give away weights. Grok 4.5 landed at nearly the same coordinates — coding, agentic tasks, app generation — undercutting on price and turning the whole category into a race for the cheapest reliable unit of finished work.
What's really being contested here isn't benchmark scores. It's ownership of the loop between what a person wants and what actually gets done — the permissions, the memory, the tool orchestration, the audit trail. A wrong sentence from a chatbot is a minor annoyance. A long-running agent that mishandles your CRM or your spreadsheets is an incident report. OpenAI's own admission that roughly 30% of tasks in the SWE-Bench Pro coding benchmark are broken — badly specified prompts, overly strict grading — only underlines how shaky the measurement tools still are for judging systems this complex.
So the frontier is quietly redefining itself around systems design instead of raw IQ: latency, token efficiency, memory management, safe rollback. The lab that wins might not have the smartest model on a leaderboard. It'll be the one whose model best schedules intelligence across tools, time, and people without anyone noticing the seams.
My take — AI-written commentary, not fact-checked reporting
I'll believe the
Read more about this at: Substack
Related stories
LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Last Week in AI · 1 month ago ·
44
AI #177 Part 1: Tip of the Iceberg
Zvi (Don't Worry About the Vase) · 1 month ago ·
32