Scores the imagined endpoint.
A useful world model can
look broken.
QuestionDoes a low planning score diagnose the learned model, or the interface wrapped around it?
MethodHold the offline-trained model fixed. Change the scoring time, replanning interval, proposal, and actuator.
ResultAligned costs repair standard tasks. Deceptive navigation still needs locally executable waypoints.
BoundaryThe study diagnoses interfaces. It does not propose a universal planner or claim that running costs are new.
A benchmark score belongs to the whole chain.
Select a link to see what the experiments changed and what remained entangled.
Plan H steps. Execute K. Then look again.
Terminal@H judges the end of an imagined sequence. A receding-horizon controller executes only K chunks before new feedback changes the plan.
No interpolation. Every selection is one measured Table 4 cell with N=200 starts.
Scores progress across the imagined rollout.
At K=3, PushT changes by +69.5 percentage points when only the scoring rule changes.
Same start. Same goal. Different score.
These clips explain one matched PushT episode. They do not establish the aggregate result.
The paper aggregates N=200 starts per standard-task cell. This retained display run contains 12 matched starts. The three clips above show one of them and use a reach-once tolerance within the 75-step budget.
The direct route points into a wall.
TwoRoom Far Door forces the first useful move away from the final goal. Choose an interface to see what path it can propose.
Watch the waypoint redirect the first move.
This pair shows one held-out-fold display episode. The 92.7% result is the five-fold mean over the paper's evaluation, not the outcome of one clip.
Ninety paths shown. Four hundred stored.
The quicklook samples 90 successful expert trajectories for legibility. The stored TwoRoom Far Door dataset contains 400 successful episodes. Each valid route climbs to the far doorway, crosses the wall, then returns toward the high final goal.

A diagnostic, not a universal planner.
- Running costs and receding-horizon MPC are established controls.
- Waypoint retrieval is a repair and a strong control, not a final planning algorithm.
- Nearest retrieval remains the strongest main held-out waypoint selector.
- The graph-value ablation uses a separate matched protocol and should be compared only within that block.
- PushT Gate has a reset-state fidelity artifact and remains an appendix sanity check.