AI Video Models Fail Their Own Physics Exam

Illustration for AI Video Models Fail Their Own Physics Exam

AI video models still cannot reliably obey the laws of physics, according to a benchmark published by researchers on arXiv on Tuesday.

Key Takeaways

  • Eight video generation models were tested across 1,280 videos and 40 controlled physics tasks
  • The best-performing model scored 57.76 out of 100 on physical consistency
  • The benchmark evaluates outputs against measurable physical relationships rather than human judgment or reference footage
  • Models that handled basic mechanics reasonably well still struggled with fluids and thermal changes

The study, titled “World Models’ Last Exam in Physics,” tested eight video generation models across 1,280 videos and found the best performer scored just 57.76 out of 100 on physical consistency. The paper covers 40 controlled tasks spanning mechanics, optics, fluids, thermal behavior, electromagnetism and surface tension.

AI video models, often called world models, generate video sequences meant to simulate how real environments behave over time.

Researchers increasingly want to use them for planning and prediction in robotics and other embodied AI systems, where a machine needs to anticipate how objects will move before acting.

The benchmark’s authors built an evaluator that pairs a starting image and a prompt with predefined physical criteria, then scores the output against measurable physical relationships rather than relying on human judgment or reference footage. That approach let the team test observable cause-and-effect relationships, like whether a dropped object accelerates correctly or whether liquid behaves consistently when poured.

Why A 57.76 Score Is A Bigger Problem Than It Sounds

A score near 58 out of 100 means the strongest model among eight tested got physical interactions right slightly more than half the time.

That gap matters because video world models are being pitched as tools for robotics planning, where a model’s internal simulation guides a physical action.

Also Read: Aleph Alpha Releases Kolibri, A Sovereign European Open-Weight Model

If a model consistently misjudges how fast something falls or how fluid disperses, any system built on top of it inherits the same errors. The researchers found the evaluator agreed with human judgments more closely than a vision-language model baseline, both in ranking videos within a single task and in head-to-head comparisons.

From Visual Realism To Physical Truth

Earlier benchmarks for video generation models mostly judged visual quality, how sharp, coherent or aesthetically convincing a clip looked, rather than whether the depicted events obeyed physical law.

This benchmark instead isolates physics as its own measurable category, splitting it into domains like thermal and phase-change phenomena where most prior evaluation tools had no dedicated test at all.

The researchers validated their measurement module on synthetic videos with known physical outcomes before applying it to the eight commercial-grade models, a step meant to confirm the scoring system itself was trustworthy before using it to judge the AI.

The gap between a 100-point physics benchmark and real-world deployment remains wide. The researchers flagged “substantial variation” across the 40 tasks, meaning models that handled basic mechanics reasonably well still stumbled on fluids and thermal changes.

Future versions of the benchmark are likely to expand task coverage as more labs release video generation systems built for robotics and simulation work.

Read Next: Two Words Nearly Double A Language Model’s Math Score

Similar Posts