PAWBench Tests Video World Models
- •PAWBench evaluates video generators as stochastic samplers for probabilistically aligned world modeling
- •PAWEval converts repeated video rollouts into empirical distributions over possible physical behaviors
- •Across 50 scenarios and eleven current systems, no model consistently matched reference probabilities
Yuandong Pu, Le Zhuo, Sayak Paul and co-authors published PAWBench on Aug 27 to measure whether video generation systems can act as probabilistically aligned world models. The paper was submitted to Hugging Face Papers by Sayak Paul on Aug 28 and ranked as the #2 Paper of the day, with 72 upvotes listed on the page.
The researchers argue that many physical processes can unfold in more than one valid way, so a world model should reproduce both a plausible single trajectory and the distribution of possible behaviors from the same initial observation and action. They define that distribution-level requirement as probabilistic alignment (matching the likelihood of possible outcomes), and say existing evaluations largely judge individual-video plausibility instead of testing repeated generations against a correct distribution.
PAWBench evaluates video generators as stochastic samplers (systems that produce varied possible outputs) of world dynamics. Its companion protocol, PAWEval, converts repeated video rollouts into empirical distributions over possible physical behaviors, allowing researchers to compare generated outcomes with reference behavior probabilities.
Across 50 scenarios and eleven current systems, the authors report that no model consistently matched the reference probabilities while also recovering the range of valid behaviors. The paper also tests whether language prompts, initial noise sampling or model training can reshape a model's predictive distribution, and positions PAWBench as a foundation for future work on probabilistically aligned world modeling.