TLDRocket
Sign in

The first GPT-6 model

Ben's Bites Covered by 20 sources

Astra is out, and it chews through tokens fast. It’s powerful, but still feels spiky, not polished.

Based on reporting by Ben's Bites — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ben’s Bites says Astra is here, and the first impression is messy in the interesting way. The model can do a lot, but it can also burn through tokens so fast that one tester spent 4 billion tokens over a weekend and “built…nothing really.” That same tester also blew through banked resets and came away annoyed, which is a pretty blunt summary of what it feels like to wrestle with a hot new model before you’ve learned its habits.

The big claim is that Astra is the start of the GPT-6 family. It is said to beat ARC-AGI-3, post the best score on Zapier’s AutomationBench, and cost the same as Fable 5.1. But the headline here is not elegance. It is volatility. One person called it a “show horse, not a workhorse,” and that seems to be the basic tension: dazzling in demos, uneven in everyday use.

Still, the demos are doing the kind of damage demos do. People are using Astra to rebuild Manhattan in Unreal Engine, generate interfaces on its own, 3D print a custom shower-drain part, identify sounds from spectrograms, and model real products in Blender and similar tools. It can also draw a portrait in Canva in the way a person would use a computer, just much better, and its computer use is fast enough to play piano.

There are practical changes too. In Codex, Astra can now keep going without waiting for answers when the next step does not need them, then fold your reply back in without losing the thread. OpenAI says that is enough to hit its automated research intern goal: agents handling tasks that would take a skilled person several days, with humans setting direction and judging the output. The next goal is an automated AI researcher by March 2028.

Elsewhere, Anthropic is testing plugins for Claude Code so it can change the interface, log actions, or limit what agents are allowed to do, though nothing has shipped yet. And Inworld has built Realtime TTS-2 for consumer apps that need text-based voice design, sub-100ms latency at P99, and one voice identity across more than 200 languages. The week’s subtext is pretty clear: the agents are getting louder, faster, and more capable. They are not getting easier to live with.

My take — AI-written commentary, not fact-checked reporting

This is the real split in AI right now: the models are getting good enough to impress everyone and annoying enough to still need babysitting. That is why the “show horse, not a workhorse” line lands so hard. The industry keeps selling magic, while users keep paying in tokens and patience.

Read more about this at: Ben's Bites

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.