TLDRocket
Sign in

The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack

Substack Jesus Rodriguez Covered by 7 sources

OpenAI, Meta, and xAI all dropped new AI models and tools last week — GPT-5.6, Grok 4.5, Muse Spark 1.1, plus agent apps. The chatbot is turning into a full work engine, not just a chat window.

Based on reporting by Substack, Jesus Rodriguez — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Last week felt less like a model release cycle and more like an infrastructure land grab. OpenAI shipped GPT-5.6 split into three variants — Sol, Terra, and Luna — each tuned for a different intelligence-per-dollar tradeoff, alongside a feature that lets the model write its own coordination programs to juggle tools and subagents. That's not autocomplete anymore. That's something closer to a scheduler.

The interface layer moved just as fast. GPT-Live ditches the old turn-based voice assistant model for full-duplex audio that can listen and talk at once, interrupt, or go quiet while reasoning runs in the background. ChatGPT Work, meanwhile, is built to sit inside a project for hours, touching files and apps and spitting out finished slide decks or spreadsheets instead of just answers. The product being sold has quietly changed from a reply to a completed artifact.

Meta and xAI answered in kind. Muse Spark 1.1 pairs a million-token context window with computer-use skills and a genuinely clever trick: it compacts long sessions without losing the state it'll need later, and it decides on the fly whether to script an action or just click through an interface itself. Meta also launched a paid Model API, a clear signal it wants to sell metered intelligence, not just give away weights. Grok 4.5 landed at nearly the same coordinates — coding, agentic tasks, app generation — undercutting on price and turning the whole category into a race for the cheapest reliable unit of finished work.

What's really being contested here isn't benchmark scores. It's ownership of the loop between what a person wants and what actually gets done — the permissions, the memory, the tool orchestration, the audit trail. A wrong sentence from a chatbot is a minor annoyance. A long-running agent that mishandles your CRM or your spreadsheets is an incident report. OpenAI's own admission that roughly 30% of tasks in the SWE-Bench Pro coding benchmark are broken — badly specified prompts, overly strict grading — only underlines how shaky the measurement tools still are for judging systems this complex.

So the frontier is quietly redefining itself around systems design instead of raw IQ: latency, token efficiency, memory management, safe rollback. The lab that wins might not have the smartest model on a leaderboard. It'll be the one whose model best schedules intelligence across tools, time, and people without anyone noticing the seams.

My take — AI-written commentary, not fact-checked reporting

I'll believe the

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.