TLDRocket
Sign in

Genie 3: A new frontier for world models

Google DeepMind

Google DeepMind just showed off Genie 3, an AI that builds entire interactive 3D worlds from a text prompt, in real time at 24fps. You can walk around inside it, and it remembers what's behind you for up to a minute.

Based on reporting by Google DeepMind — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google DeepMind has a new toy, and it's a strange one: type a sentence, and out pops a navigable world. Genie 3, unveiled August 5, generates interactive environments at 720p and 24 frames per second, and unlike a video, you can actually steer through it, turn corners, revisit spots, watch a scene hold together for a few minutes at a stretch.

That consistency is the real headline here. Earlier world models, including DeepMind's own Genie 1 and Genie 2 from last year, could conjure environments but struggled to keep them coherent as you moved through them. Genie 3 apparently holds visual memory roughly a minute deep, so if you walk away from a building and come back, the trees are still where you left them. No explicit 3D map underneath, no NeRF-style scaffolding — the whole thing is generated frame by frame, reacting to your inputs on the fly. DeepMind calls this an emergent property, which is a polite way of saying even they're a little surprised it works.

Beyond just walking around, Genie 3 supports what the team calls promptable world events: mid-simulation, you can type something like "make it rain" or drop a new character into the scene, and the world adjusts. That's less a parlor trick than a training tool. DeepMind ran its SIMA agent through Genie-generated worlds, handing it goals and watching it navigate toward them, with the model itself blind to what the agent was actually trying to achieve. It's a cheap, infinite way to manufacture practice environments for embodied AI, robots and virtual agents included, without building a single physical set.

The limitations are real and DeepMind doesn't hide them. Real-world locations come out geographically fuzzy. Text in scenes only renders cleanly if it was already in the prompt. Multiple independent agents interacting with each other remains mostly unsolved. And sessions cap out at a few minutes, not hours. For now, Genie 3 is a research preview reserved for a small group of academics and creators, which suggests DeepMind knows this thing needs guardrails before it goes anywhere near the general public.

Still, the pitch is bigger than a demo. DeepMind frames this explicitly as a stepping stone toward AGI, the idea being that if you can generate unlimited, varied, physically plausible worlds on demand, you can train agents in something close to an infinite curriculum. Whether that pans out or just produces very convincing screensavers is the open question.

My take — AI-written commentary, not fact-checked reporting

I'll believe the AGI-stepping-stone framing when Genie 3 can hold a world together for longer than a coffee break, not just a few minutes — right now this reads more like an extremely fancy procedural sandbox than a training ground for anything general. That said, gating it to academics and creators instead of shipping it straight to the public is the correct call, and it's refreshing to see a lab actually act on that instinct instead of just saying it in a blog post.

Read more about this at: Google DeepMind

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.