Dynamic Scene Reconstruction
Dynamic scene reconstruction estimates persistent identities, canonical geometry, and time-varying poses in a shared 3D coordinate system.
Paper
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.
CLEVRER & Physion++
PhysMind recovers a temporally consistent dynamic scene from each video, then fits and executes the recovered dynamics. Given a question, it inspects, continues, or edits the world.
PhysMind proceeds through two sequential stages: dynamic scene reconstruction recovers temporally consistent geometry and motion, after which executable physical world modeling fits and executes the recovered dynamics.
Dynamic scene reconstruction estimates persistent identities, canonical geometry, and time-varying poses in a shared 3D coordinate system.
PhysMind converts the dynamic scene into an executable 3D world by fitting continuous-time analytic dynamics to the recovered trajectories and meshes.
Experiments
PhysMind attains the highest overall accuracy among the evaluated methods on CLEVRER and Physion++.
| Method | Explanatory | Predictive | Counterfactual | Overall | ||||
|---|---|---|---|---|---|---|---|---|
| per ques. | per opt. | per ques. | per opt. | per ques. | per opt. | per ques. | per opt. | |
| Baselines | ||||||||
| Random | 7.19 | 49.33 | 25.40 | 49.71 | 9.75 | 50.02 | 11.26 | 49.69 |
| Blind Gemini-3-Flash | 27.00 | 64.64 | 24.68 | 48.92 | 10.07 | 50.13 | 19.21 | 56.31 |
| Foundation VLMs | ||||||||
| Qwen3-VL-235B-A22B | 53.19 | 75.84 | 40.40 | 62.12 | 22.17 | 61.79 | 37.52 | 67.92 |
| GLM-4.6V | 39.33 | 73.75 | 63.64 | 66.74 | 24.31 | 58.70 | 36.68 | 66.02 |
| Gemini-3-Flash | 50.32 | 74.66 | 42.71 | 53.46 | 16.63 | 58.38 | 34.32 | 64.97 |
| Gemini-3.1-Pro | 60.90 | 79.17 | 61.76 | 71.28 | 25.37 | 63.50 | 45.47 | 71.06 |
| GPT-4o | 36.76 | 70.48 | 35.64 | 54.98 | 15.57 | 56.54 | 27.29 | 62.44 |
| GPT-5.5 | 77.21 | 89.55 | 85.71 | 90.69 | 51.33 | 76.74 | 67.24 | 83.66 |
| Training-Based Methods | ||||||||
| VideoRFT | 16.66 | 61.98 | 39.25 | 47.40 | 15.35 | 54.98 | 19.74 | 57.28 |
| Video-R1 | 32.79 | 71.14 | 41.56 | 46.25 | 19.62 | 59.10 | 28.43 | 63.08 |
| VideoThinker-R1 | 17.53 | 63.60 | 41.56 | 42.50 | 15.14 | 53.71 | 20.37 | 56.92 |
| Chain-of-Frames | 44.71 | 77.26 | 84.42 | 87.88 | 33.48 | 70.84 | 46.21 | 75.29 |
| Training-Free Methods | ||||||||
| VideoAgent | 34.42 | 61.55 | 14.29 | 48.56 | 7.57 | 51.61 | 19.39 | 55.63 |
| STAR | 55.58 | 78.62 | 24.24 | 57.14 | 7.57 | 53.47 | 29.46 | 64.75 |
| PhysMind | 76.97 | 87.38 | 66.96 | 82.32 | 70.58 | 88.08 | 72.55 | 87.22 |
| Method | Fric. Plat. | Fric. Coll. | Bounce Plat. | Bounce Wall | Mass Coll. | Overall |
|---|---|---|---|---|---|---|
| Baselines | ||||||
| Random | 48.44 | 45.31 | 53.12 | 51.04 | 43.75 | 48.96 |
| Blind Gemini-3-Flash | 46.88 | 50.00 | 47.92 | 52.08 | 51.56 | 49.74 |
| Foundation VLMs | ||||||
| Qwen3-VL-235B-A22B | 48.44 | 43.75 | 56.25 | 50.00 | 42.19 | 48.96 |
| Gemini-3-Flash | 56.25 | 46.88 | 52.08 | 50.00 | 53.13 | 51.56 |
| GPT-5.5 | 51.56 | 65.62 | 55.21 | 58.33 | 60.94 | 58.07 |
| Training-Based Methods | ||||||
| Video-R1 | 54.69 | 53.13 | 51.04 | 45.83 | 53.13 | 51.04 |
| Chain-of-Frames | 48.44 | 51.56 | 50.00 | 51.04 | 50.00 | 50.26 |
| Training-Free Methods | ||||||
| PhysMind | 54.69 | 64.06 | 57.29 | 59.38 | 64.06 | 59.64 |
PhysMind costs $0.01654 per question. It improves overall accuracy over GPT-5.5 by 5.31 points, while GPT-5.5 costs 6.30 times as much.
Replacing analytic identification with fixed-step numerical integration produces the largest overall loss. Execution is specifically responsible for the counterfactual improvement.
The four world-construction categories account for 92.31% of errors. Errors after valid artifacts are uncommon.