Selected work KV / TRA 3D 2D clip page
ICLR 2026 · 2nd Workshop on World Models · accepted · corrected July 2026 · 3D companion

Pull time
backward.

The clip page watches 48 recorded frames through a window. This page hands you the world: a domino run you can orbit while you drag time in both directions. Forward, the motion spends its energy and goes still. Backward, a settled pile climbs back onto its edges and a resting ball gathers speed for nothing — dissipation running in reverse, which is the entire paper.

The world below is a re-simulated lookalike built for intuition — the repo holds rendered clips, not object poses. The recorded clips and their exact losses stay in the clip lab, and the paper’s corrections stay in the evidence receipt. Corrected paper ↗

Rehearsal · re-simulated in your browser — not the recorded clip
F 01 / 48 ORDER FORWARD CTX 8
Reversal roomorbit: drag · zoom: scroll

The run is playing. Watch where the energy goes.

9 dominoes · 48 frames @ 12 fps · hand-authored motion

Time transport← → scrub · space play · d flip
PLAYED 01 / 48 · SOURCE 01 / 48or drag the floor strip
Motion · proxy
Pile height · proxy

Geometric proxies of the staged motion — watch them refill when time runs backward. Not model outputs.

Context windowthe experiment’s shape
GIVEN 08 · PREDICTED 40context frames N

CONTEXT 14 · the awkward cell. In the corrected randomized-domino rerun, V-JEPA2’s separation flips sign exactly here — reversed becomes the easier prediction. The red band on the rail pin marks this boundary. Exact values sit in the clip lab.

Model lensqualitative only

Arrow testscore 0 / 0

A hidden direction plays in the world. Orbit around it — being able to inspect from any angle and still be unsure is the point — then call it.

CORRECT

Free play, world focused: ← → scrub · space play/pause · D flip direction · 1–4 cameras

Provenance
What is exact
Every loss value, context sweep, rerun mean, and withdrawal on this page is quoted from the corrected paper and the three checked-in 48-frame PyBullet sequences. The clip lab plays the real mp4 files and attaches the real numbers to them, per scene, model, and context.
What is staged
The 3D world above. The repository contains rendered clips only — no per-frame poses — so the dominoes you scrub are hand-authored keyframe motion: a different seed, solver, and pile than any recorded run. It demonstrates the intuition and produces no loss values. Its energy bars are geometric proxies of that staged motion, not model outputs.

The recorded sequences,
with their losses attached.

Six direction clips from three checked-in 48-frame PyBullet sequences. Each loss value below was read from video_000 for the selected scene and context, on the footage playing beside it. This lab is the evidence; the 3D world above is not.

Scene
Model
Context frames
bouncing-forward.mp448 frames · 12 fps · PyBullet
bouncing-reverse.mp448 frames · 12 fps · PyBullet
V-JEPA2 · BOUNCING · CONTEXT 8 +0.0000%

L forward · exact
L reverse · exact
SceneBOUNCING
Condition

Values quoted from the corrected paper’s per-clip evaluation of the checked-in sequences. They describe these mp4s — not the 3D world above, not the 160-video randomized-domino rerun, and not the 220-video continuous sweep.

TRA across contextexact clip · same model · click a column

TRA = (L reverse − L forward) / L forward, on this exact clip. Positive: the reversed order costs more to predict.

Know what each number can prove.

This page mixes a playable lookalike with checked-in measurements, so the boundaries are drawn explicitly. Five scopes, no merging.

00 · REHEARSAL

The 3D world is not evidence

Hand-authored motion re-created in your browser. The repo holds rendered clips only — no per-frame poses — so this is a different seed, solver, and pile. It earns no numbers.

01 · EXACT CLIP

Real sequence, real loss

Six direction clips come from three checked-in 48-frame PyBullet sequences. The lab’s loss values change with scene, model, and context, read from video_000 per selection.

02 · AGGREGATE

380 simulations

The paper evaluates 160 discrete-scene videos and 220 continuous restitution or damping videos. A clip is an example, not the claim.

03 · CORRECTED CLAIM

Two supported models

V-JEPA2 shows the expected dissipative separation for contexts 4–12 and reverses sign at context 14. VideoMAE v1 shows the inverse pattern throughout.

04 · NOT AVAILABLE

Missing artifacts stay out

The corrected rerun has no matching frames checked in here; continuous-sweep videos and per-frame loss traces are also absent. The 3D world cannot regenerate any of them.

Open aggregate evidence and correction notes
SUPPORTED

V-JEPA2

Positive dissipative separation holds across contexts 4–12 in the corrected randomized-domino rerun. The aggregate sign reverses at context 14.

SUPPORTED

VideoMAE v1

The inverse pattern survives recomputation. The corrected paper fixes the checkpoint name from VideoMAE v2 to VideoMAE v1.

WITHDRAWN

MVD

The public checkpoint had no trained decoder weights. Its result cannot support a claim about distillation.

WITHDRAWN

Hiera

The evaluator used random-mask inpainting instead of symmetric future prediction. The objective comparison is invalid.

Corrected randomized-domino rerunΔ = dissipative − low-dissipation TRA
V-JEPA2 +0.1185+0.2084+0.1708+0.1207+0.1702−0.2187
VideoMAE v1 −0.1203−0.1744−0.2375−0.2868−0.3308−0.4015

Percentage points. Exact means from the corrected 160-video discrete rerun.

Corrected dataset with 40 randomized domino seeds. V-JEPA2 separates the groups at contexts 4–12 and reverses at 14; VideoMAE v1 remains inverted.
Original-dataset TRA context sweep for V-JEPA2 EMA, V-JEPA2 without EMA, and VideoMAE v1
Original-dataset aggregate after correcting the model identity: V-JEPA2 EMA, V-JEPA2 without EMA, and VideoMAE v1. This is not the randomized-domino rerun.
Method diagram comparing forward and reversed prediction losses
The same frozen prediction pipeline receives both frame orders.

TRA measures a change in prediction loss. It does not prove that a model understands physics. The core evidence is synthetic, and the real-video results remain confounded by human action timing.