NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks
MarkTechPost Asif Razzaq
NVIDIA and partners built Physis-Lang, a physics-aware caption system that helps video models understand why scenes happen. It pushed Cosmos3-Super to the top of a physics leaderboard and beat Veo 3.1 there.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Video models can look slick and still get basic physics wrong. Butter behaves like paint. Objects ignore walls. NVIDIA, MIT and the University of Oxford are trying to fix that by pushing more of the burden onto language itself, with a framework called Physis-Lang that treats physical language as the shared layer for curation, training and inference.
The idea is simple enough: don’t just describe what a clip shows, describe why it unfolds that way. Physis-Lang adds a physics_reasoning field to each caption, spelling out the entities in a scene, the causes, interactions, governing principles, timing and effects. It also generates a scene-specific physics_negative_prompt, which names implausible outcomes and is used as negative conditioning when the model is run.
The team’s captioning loop is built to improve itself. A GPT-5.5 captioner works on a fixed 20-video development set with 273 human-verified assertions, while Gemini-3.1-Pro acts as a physics-aware critic. An evolution agent watches the claim-level failures and rewrites the prompt, then the revised prompt is checked on PhysCapBench, a benchmark of 246 videos and 3,794 human-verified assertions. Caption F1 moved from 78.64 at iteration 1 to 87.82 at iteration 9, after a rough dip to 76.28 at iteration 2.
The gains weren’t automatic. By iteration 9, the prompt had to require every visible causal step, and the frame sampling rate had to go from 2 fps to 4 fps. That tells you this is less about magical prompting and more about grinding until the captions stop hand-waving around the physics.
Physis-Lang also feeds training data selection. A GPT-5.5 diagnosis agent maps generated-video failures to categories like rigid-body motion, collision and fluid dynamics, then those gaps are matched against physics tags in a large video gallery. The final training set has 183K videos: 71K filtered from WISA-80K and 112K retrieved clips. Retrieval alone improved performance by 3.01 points on average across 3 benchmarks, and on VideoPhy-2 the chemical and thermal process categories each gained 8.00 points. On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4, with Cosmos3-Nano in second at 43.3 ± 1.5.
My take — AI-written commentary, not fact-checked reporting
This is the rare AI paper that sounds useful instead of decorative. Language keeps showing up as the cheapest way to force models to stop bluffing about the physical world, which is awkward for the people selling pure-vision magic. Turns out “describe the physics properly” is still a better plan than hoping the pixels sort it out themselves.
Read more about this at: MarkTechPost
Related stories
Into the Omniverse: How Open World Models Push the Frontier of Physical AI
NVIDIA · 1 month ago ·
24
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
NVIDIA · 2 months ago ·
26
NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets
Hugging Face · 1 year ago ·
57