Introducing Muse Image and Muse Video
Meta AI ● Covered by 3 sources
Meta's new Superintelligence Labs dropped Muse Image today and teased Muse Video, its first media-gen models. Muse Image already ranks No. 2 on human-preference leaderboards and can search, code, and fix its own mistakes mid-generation.
Based on reporting by Meta AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Meta Superintelligence Labs, the group Mark Zuckerberg's company stood up to chase frontier AI, just shipped its first media products: Muse Image, live now, and Muse Video, still in preview. Muse Image is rolling out across the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook support coming soon. Muse Video, which shares the same pretraining base and adds native audio, is headed to creators and Meta AI but has no firm release date yet.
What separates Muse Image from a typical prompt-to-picture model is that it doesn't just generate — it acts. During reinforcement learning, the model learned to write and run code so it can produce accurate plots and QR codes, and it can search the web to ground an image in real facts or visual references. Meta says this search capability specifically boosts accuracy on prompts tied to current events or real-world details, the kind of thing pure pattern-matching models tend to fumble.
The more interesting wrinkle is self-refinement, and Meta is upfront that they didn't design it on purpose. Somewhere in training, the model discovered that reflecting on its own draft and fixing it — sometimes a small local edit, sometimes scrapping the image and starting over, sometimes switching to a tool like search — simply earned higher reward. So it kept doing it. Layered on top is test-time compute scaling: give Muse Image more time to reason, call tools, and refine, and quality climbs in a roughly log-linear curve. Meta also found that just generating a pile of images and picking the best one (best-of-N) plateaus fast, while spending that same compute on deliberate reasoning keeps paying off, especially when reasoning and tool use are combined.
On the editing side, Muse Image can make precise, targeted changes across multiple turns without losing coherence, and it can compose a single image from several reference inputs — people, clothing, objects, styles, environments — mixing text and images inline in the same prompt. Meta says this combination of instruction-following, editing precision, and multi-reference composition puts Muse Image in the No. 2 spot on Arena's human-preference rankings for text-to-image, single-image editing, and multi-image editing, as of July 5, 2026. Muse Video, still early, sits at No. 3 for text-to-video on the same leaderboard, with Meta acknowledging it still lags on audio-video sync and fast-motion physics.
Every image Muse Image produces inside the Meta AI app and meta.ai carries an invisible watermark called Content Seal, designed to survive cropping, compression, resizing, and even screenshots. Meta is also previewing a standalone detection tool so people can check whether an image was actually made with Meta AI, and says video watermarking is coming next.
My take — AI-written commentary, not fact-checked reporting
Landing at No. 2 and No. 3 on the leaderboards is a perfectly respectable debut, not a coronation, and Meta seems to know it by leaning so hard on the agentic story — search, code, self-correction — rather than claiming outright supremacy. The self-refinement behavior emerging unplanned from reward signals is the genuinely interesting part here, more so than the rankings, because it suggests these systems are starting to develop their own workflows rather than just following one. The watermarking push is the right instinct, but a detection tool only matters if people actually bother to use it before sharing something.”}}
Read more about this at: Meta AI