TLDRocket
Sign in

NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

MarkTechPost Asif Razzaq

NVIDIA and partners built Physis-Lang, a physics-aware caption system that helps video models understand why scenes happen. It pushed Cosmos3-Super to the top of a physics leaderboard and beat Veo 3.1 there.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Video models can look slick and still get basic physics wrong. Butter behaves like paint. Objects ignore walls. NVIDIA, MIT and the University of Oxford are trying to fix that by pushing more of the burden onto language itself, with a framework called Physis-Lang that treats physical language as the shared layer for curation, training and inference.

The idea is simple enough: don’t just describe what a clip shows, describe why it unfolds that way. Physis-Lang adds a physics_reasoning field to each caption, spelling out the entities in a scene, the causes, interactions, governing principles, timing and effects. It also generates a scene-specific physics_negative_prompt, which names implausible outcomes and is used as negative conditioning when the model is run.

The team’s captioning loop is built to improve itself. A GPT-5.5 captioner works on a fixed 20-video development set with 273 human-verified assertions, while Gemini-3.1-Pro acts as a physics-aware critic. An evolution agent watches the claim-level failures and rewrites the prompt, then the revised prompt is checked on PhysCapBench, a benchmark of 246 videos and 3,794 human-verified assertions. Caption F1 moved from 78.64 at iteration 1 to 87.82 at iteration 9, after a rough dip to 76.28 at iteration 2.

The gains weren’t automatic. By iteration 9, the prompt had to require every visible causal step, and the frame sampling rate had to go from 2 fps to 4 fps. That tells you this is less about magical prompting and more about grinding until the captions stop hand-waving around the physics.

Physis-Lang also feeds training data selection. A GPT-5.5 diagnosis agent maps generated-video failures to categories like rigid-body motion, collision and fluid dynamics, then those gaps are matched against physics tags in a large video gallery. The final training set has 183K videos: 71K filtered from WISA-80K and 112K retrieved clips. Retrieval alone improved performance by 3.01 points on average across 3 benchmarks, and on VideoPhy-2 the chemical and thermal process categories each gained 8.00 points. On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4, with Cosmos3-Nano in second at 43.3 ± 1.5.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI paper that sounds useful instead of decorative. Language keeps showing up as the cheapest way to force models to stop bluffing about the physical world, which is awkward for the people selling pure-vision magic. Turns out “describe the physics properly” is still a better plan than hoping the pixels sort it out themselves.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.