Into the Omniverse: How Open World Models Push the Frontier of Physical AI
NVIDIA Ming-Yu Liu
NVIDIA just launched Cosmos 3, an open family of AI models that predict how the physical world behaves. It's meant to help robots, self-driving cars and cameras train and test before touching the real world.
Based on reporting by NVIDIA, Ming-Yu Liu — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA has released Cosmos 3, the newest version of its open world foundation model family, built to help physical AI systems like robots, autonomous vehicles and vision AI understand and predict how the real world actually behaves. Rather than just recognizing what's in a scene, these models are meant to simulate what happens next and help teams generate the training data they need without capturing every rare or dangerous scenario in the real world.
Cosmos 3 is built on a mixture-of-transformers architecture and comes in three sizes: Super at 64 billion parameters for high-fidelity world modeling, Nano at 16 billion for efficient reasoning and post-training, and Edge at 4 billion for on-device use on platforms like Jetson Thor. One model family can act as a vision language model, a physics-grounded simulator predicting future states, or the backbone for what NVIDIA calls world action models — collapsing what used to require separate, maintained models into a single stack.
The benchmark numbers NVIDIA is touting are notable. Cosmos 3 ranks first on Artificial Analysis for open-weights text-to-image and image-to-video generation, first on PAI-Bench for world generation, first in the image-to-video category of Physics-IQ, and first on RoboLab for robot policy. Cosmos 3 Super also tops VANTAGE-Bench for vision understanding among open models. NVIDIA released the models under the Linux Foundation's OpenMDW 1.1 license, which matters because it lets teams actually post-train on their own hardware and data — a general model has never seen a specific robot's sensors or a particular warehouse floor, and closing that gap requires real access to weights, not just an API.
Adoption already spans multiple industries. Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI are working with Cosmos in robotics; Li Auto, Xiaomi and Afari in autonomous vehicles; and Centific, Fogsphere, Linker Vision, Milestone Systems and Yuan in vision AI for industrial and smart-space applications. NVIDIA also expanded its Cosmos Coalition — a group of model builders and physical AI companies contributing research and evaluation methods — into Japan, where manufacturing and robotics firms are lining up to build open world models for factories, logistics, agriculture, construction, healthcare and transportation.
NVIDIA frames all of this alongside its broader physical AI stack — Isaac GR00T for robotics, Alpamayo for autonomous vehicles, Metropolis for vision AI — plus Omniverse and OpenUSD tools for building the simulation environments these models need. The pitch is straightforward: open weights plus real post-training access is what turns a generic world model into something a specific robot or vehicle can actually use.
My take — AI-written commentary, not fact-checked reporting
The interesting move here isn't the benchmark sweep, it's the licensing. Slapping
Read more about this at: NVIDIA
Related stories
NVIDIA Releases Cosmos 3 Edge: A 4B-Parameter Open World Model That Reasons and Generates Robot Actions On-Device
MarkTechPost · 2 months ago ·
8
Introducing Cosmos 3 Edge
Hugging Face · 2 months ago ·
45