Video world models can render convincing clips that still break the laws of physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT, and the University of Oxford argues that the fix can come from language itself—not from extra visual, latent, or numerical signals.
What Is Physis-Lang?
Their framework, Physis-Lang, treats physical language as a shared, optimizable representation. The same text drives data curation, model training, and inference.
On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4, while the Cosmos3-Nano version ranks second at 43.3 ± 1.5. The results position NVIDIA's Cosmos 3 family ahead of Veo 3.1 on physics-specific benchmarks.
What Problem Does Physis-Lang Solve?
Conventional captions describe what appears in a scene, not why or how it unfolds. As a result, video world models trained on such captions can produce visually plausible outputs that violate basic physical constraints. Physis-Lang addresses this gap by embedding physics reasoning directly into the language used throughout the pipeline.
Rather than adding separate physics engines or simulation signals, the framework makes the caption itself a physics-aware representation. This single change propagates through data curation, training, and inference, creating a self-evolving loop where physical language improves the model and the model, in turn, refines the language.
Why It Matters in 2026
As video generation moves from novelty to production infrastructure, physical fidelity has become a key differentiator. Applications in robotics, autonomous systems, and physical AI depend on world models that respect real-world dynamics, not just visual plausibility. By topping the Physics-IQ Verified leaderboard, Physis-Lang signals that language-based physics reasoning may be a more scalable path than adding increasingly large visual or numerical signal streams.
Key Takeaways
- Physis-Lang is an open, self-evolving framework that adds physics reasoning to video captions.
- Cosmos3-Super + Physis-Lang leads the Physics-IQ Verified leaderboard at 48.2 ± 1.4.
- Cosmos3-Nano + Physis-Lang ranks second at 43.3 ± 1.5.
- The approach uses a single language representation across data curation, training, and inference.
- The work is a collaboration between NVIDIA, MIT, and the University of Oxford.
via MarkTechPost
