Counterfactual World Models · Motion Reasoning · Physical Adjudication

LPA-CWM

A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models
Kunwei Wu1,* Xiang Liu2,* Guocai Yao3 Junming Chen4 Zhikang Chen5 Min Zhang6 Pengwei Wang3 Sen Cui2,†
1National University of Singapore  ·  2Tsinghua University  ·  3Beijing Academy of Artificial Intelligence  ·  4Facebook  ·  5University of Oxford  ·  6East China Normal University
*Equal contribution    Project Leader & Corresponding Author
Explore
01 · Overview

Learn which counterfactual motion evidence to trust.

Uniform CWM treats responses from different target-frame masks equally. LPA-CWM instead learns candidate reliability and uses the weighted evidence to recover more coherent motion.
Figure 1: overview of LPA-CWM and the CMC benchmark
Figure 1 · An introduction of our work. LPA identifies reliable counterfactual motion evidence, improving trajectory coherence under drift, competing motion, and appearance changes; CMC evaluates localization, completeness, visibility, and continuity.
Frozen backboneCWM predictor and intervention generator remain fixed.
Lightweight LPAOnly a 3.0M-parameter adjudicator is trained.
Reliability learningCandidate responses receive learned relative weights.
CMC benchmarkLocalization, completeness, visibility, and continuity.
02 · Method

Analytic and learned adjudication over frozen CWM candidates.

APA-CWM provides a training-free physical reference, while LPA-CWM learns relative candidate reliability from dense MOVi-F trajectories and applies the resulting weights before localized decoding and one CWM re-evaluation.
Figure 2: overview of APA-CWM and LPA-CWM
Figure 2 · Overview of APA-CWM and LPA-CWM. A frozen CWM produces candidate responses. APA uses NSG-based neighborhood scores, whereas LPA encodes visual context, response structure, frame context, and candidate statistics with set-based reasoning.
APA-CWM

Analytic Physical Adjudicator

Training-free candidate weighting from a frozen ADM / NSG physical-consistency cue.

LPA-CWM

Learned Physical Adjudicator

Candidate-specific tokens are jointly compared by a Set Transformer and supervised using MOVi-F endpoint errors.

Inference

Weighted response → refinement

Weighted responses are locally decoded, followed by one paired CWM re-evaluation to obtain refined motion.

03 · CMC Benchmark

Completeness-aware Motion Correspondence.

Coming soonTODO · benchmark visualization / metric design / figure assets
04 · Quantitative Results

DAVIS, Kinetics, and TAP-Vid First.

Coming soonTODO · main tables / plots / interactive result views
05 · Qualitative Results

Motion correspondence in the wild.

Coming soonTODO · videos / tracking comparisons / manipulation examples
06 · Ablation Studies

What makes learned adjudication work?

Coming soonTODO · mask count / re-evaluation / architecture / objective / weighting behavior
07 · Models & Data

Code, checkpoints, and reproducibility assets.

Coming soonTODO · GitHub repository / trained weights / configs / processed data
08 · Citation

BibTeX and acknowledgements.

Coming soonTODO · arXiv BibTeX / acknowledgements / contact links