Diffusion-forcing world model rollouts. Each sample is ~2.7 s (16 raw frames @ 6 fps). Side-by-side videos show ground truth on the left, model prediction on the right. First half of each clip = history (GT), second half = autoregressively sampled future. Click a sample header to expand its videos (default-collapsed to keep the page light).
Click a run name to jump to its rollout videos. ✓ = feature on, ✗ = off. Lower val_loss is better.
| run | fusion | views | tactile | shift16 | delta-ref | cam-pose | gate | val_loss |
|---|---|---|---|---|---|---|---|---|
| 12v_vision_only_val_0.0093 | — | — | — | — | — | — | — | — |
Latent-space MSE of predicted future vs ground truth. Contact/no-contact split uses tactile latent-to-reference energy (threshold 0.05).
| run | view MSE | TL MSE | TR MSE | tactile contact MSE | tactile no-contact MSE |
|---|---|---|---|---|---|
| 12v_vision_only_val_0.0093 | 1.737214 | 5.588249 | 5.637631 | 5.582695 | 5.658344 |