The 3D world is not evidence
Hand-authored motion re-created in your browser. The repo holds rendered clips only — no per-frame poses — so this is a different seed, solver, and pile. It earns no numbers.
The clip page watches 48 recorded frames through a window. This page hands you the world: a domino run you can orbit while you drag time in both directions. Forward, the motion spends its energy and goes still. Backward, a settled pile climbs back onto its edges and a resting ball gathers speed for nothing — dissipation running in reverse, which is the entire paper.
The world below is a re-simulated lookalike built for intuition — the repo holds rendered clips, not object poses. The recorded clips and their exact losses stay in the clip lab, and the paper’s corrections stay in the evidence receipt. Corrected paper ↗
9 dominoes · 48 frames @ 12 fps · hand-authored motion
Geometric proxies of the staged motion — watch them refill when time runs backward. Not model outputs.
CONTEXT 14 · the awkward cell. In the corrected randomized-domino rerun, V-JEPA2’s separation flips sign exactly here — reversed becomes the easier prediction. The red band on the rail pin marks this boundary. Exact values sit in the clip lab.
A hidden direction plays in the world. Orbit around it — being able to inspect from any angle and still be unsure is the point — then call it.
Free play, world focused: ← → scrub · space play/pause · D flip direction · 1–4 cameras
Six direction clips from three checked-in 48-frame PyBullet sequences. Each loss value below was read from video_000 for the selected scene and context, on the footage playing beside it. This lab is the evidence; the 3D world above is not.
Values quoted from the corrected paper’s per-clip evaluation of the checked-in sequences. They describe these mp4s — not the 3D world above, not the 160-video randomized-domino rerun, and not the 220-video continuous sweep.
TRA = (L reverse − L forward) / L forward, on this exact clip. Positive: the reversed order costs more to predict.
This page mixes a playable lookalike with checked-in measurements, so the boundaries are drawn explicitly. Five scopes, no merging.
Hand-authored motion re-created in your browser. The repo holds rendered clips only — no per-frame poses — so this is a different seed, solver, and pile. It earns no numbers.
Six direction clips come from three checked-in 48-frame PyBullet sequences. The lab’s loss values change with scene, model, and context, read from video_000 per selection.
The paper evaluates 160 discrete-scene videos and 220 continuous restitution or damping videos. A clip is an example, not the claim.
V-JEPA2 shows the expected dissipative separation for contexts 4–12 and reverses sign at context 14. VideoMAE v1 shows the inverse pattern throughout.
The corrected rerun has no matching frames checked in here; continuous-sweep videos and per-frame loss traces are also absent. The 3D world cannot regenerate any of them.
Positive dissipative separation holds across contexts 4–12 in the corrected randomized-domino rerun. The aggregate sign reverses at context 14.
The inverse pattern survives recomputation. The corrected paper fixes the checkpoint name from VideoMAE v2 to VideoMAE v1.
The public checkpoint had no trained decoder weights. Its result cannot support a claim about distillation.
The evaluator used random-mask inpainting instead of symmetric future prediction. The objective comparison is invalid.
Percentage points. Exact means from the corrected 160-video discrete rerun.


TRA measures a change in prediction loss. It does not prove that a model understands physics. The core evidence is synthetic, and the real-video results remain confounded by human action timing.