# `media/` provenance

Generated by `scripts/website/build_website_videos.py` -- do not hand-edit.

Every clip is a real pipeline artifact. Sources are the episodes and
evaluation samples already frozen by the paper's figure provenance records
(`overleaf/figures/qualitative_checkpoint_eval_comparison.json`,
`overleaf/assets/teaser_assets/README.md`,
`overleaf/assets/method_overview_asseets/MANIFEST.md`,
`overleaf/assets/action_flow_assets/provenance.json`,
`robolab/annotations/manifest.json`), so a clip on this page and the same
clip in the paper are the same bytes upstream of the crop.

One exception, and it is deliberate. The two RoboLab rows are re-renders
rather than the rated corpus bytes: the rating render draws the
conditioning tracks into the prediction pane as cyan dots, a debug aid
that on this page reads as an artefact of the prediction. They were rolled
out again at the same LoRA step 6000, the same episode, the same
conditioning pool and the same seed with `--no-draw-tracks`. Both come
back at the same length as the rated clip (109 and 277 frames), and the
prediction pane agrees with the rated one to within re-encode noise
outside the pixels the overlay used to cover.

The method and sampling sections play what those two figures could only
sample: same episodes, same window starts, same tracks, same selection
rules and the same drawn counts (the deployment lane draws 98 of its 365
FK tracks, the Deform360 tile 29 embodiment and 37 object -- the numbers
the shipped tiles report), with the frame index running instead of fixed.

Each slot ships a `posters/<slot>.jpg` cut from its own clip, so the still a
browser shows before playback is always a frame of what it is about to play.

## Figure 1 band 1 - heterogeneous sources

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/teaser_source_deform360.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, Deform360 | centre-crop to 4:3, resize to 640x480 |
| `media/teaser_source_droid.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, DROID | centre-crop to 4:3, resize to 640x480 |
| `media/teaser_source_egodex.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, EgoDex | centre-crop to 4:3, resize to 640x480 |
| `media/teaser_source_xvla_yam.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, XVLA-Soft-Fold | centre-crop to 4:3, resize to 640x480 |

## Figure 1 band 2 - shared action flow

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/teaser_flow_deform360.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, Deform360 | progressive per-track trail, then the same 4:3 crop |
| `media/teaser_flow_droid.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, DROID | progressive per-track trail, then the same 4:3 crop |
| `media/teaser_flow_egodex.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, EgoDex | progressive per-track trail, then the same 4:3 crop |
| `media/teaser_flow_xvla_yam.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, XVLA-Soft-Fold | progressive per-track trail, then the same 4:3 crop |

## Figure 1 band 3a - SIMULATE

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/teaser_simulate_deform360.mp4` | 640x480 | 81 | 16 | eval bundle predicted.mp4, deform360 held-out sample | centre-crop to 4:3, resize to 640x480 |
| `media/teaser_simulate_droid.mp4` | 640x480 | 81 | 16 | eval bundle predicted.mp4, droid_raw held-out sample | centre-crop to 4:3, resize to 640x480 |

## Figure 1 band 3b - EVALUATE

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/teaser_eval_failure_robolab.mp4` | 640x480 | 277 | 16 | rated RoboLab episode, failure row, re-rendered without the conditioning overlay -- RoboLab reference pane | same pane split, then centre-crop to 4:3 |
| `media/teaser_eval_failure_sim.mp4` | 640x480 | 277 | 16 | rated RoboLab episode, failure row, re-rendered without the conditioning overlay -- our neural simulator pane | same pane split, then centre-crop to 4:3 |
| `media/teaser_eval_success_robolab.mp4` | 640x480 | 109 | 16 | rated RoboLab episode, success row, re-rendered without the conditioning overlay -- RoboLab reference pane | same pane split, then centre-crop to 4:3 |
| `media/teaser_eval_success_sim.mp4` | 640x480 | 109 | 16 | rated RoboLab episode, success row, re-rendered without the conditioning overlay -- our neural simulator pane | same pane split, then centre-crop to 4:3 |

## Figure 1 band 3c - CONTROL

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/teaser_control_human.mp4` | 556x418 | 107 | 15 | GEAR-YAM/wam_rollout/pipe_play, human demonstration with its derived object flow | crop 62 rows off the bottom, then centre-crop to 4:3 (557x418, matching the teaser stills) |
| `media/teaser_control_robot.mp4` | 556x418 | 68 | 15 | GEAR-YAM/wam_rollout/pipe_play, physical YAM execution driven by that flow | crop 62 rows off the bottom, then centre-crop to 4:3 (557x418, matching the teaser stills) |

## Figure 3 - qualitative rollouts

Three held-out evaluation samples per dataset, four methods each, curated
from a seeded draw of ten and named by sample id in
`extract_website_sources.py`. Sample 1 is the id
`qualitative_checkpoint_eval_comparison.json` freezes, so the lead of
XVLA-Soft-Fold, Deform360 and DROID is the clip the paper figure shows;
ABC-130k has no figure counterpart and leads with a curated sample.

Each sample ships four slots -- `<row>__s<NN>__{wan_move,cosmos2p5,ours,gt}`
-- resampled into that sample's own ground-truth cell shape. The shape is
tracked per sample, not per dataset: ABC-130k renders 640, 768 and 848
wide across its evaluation set, so a per-row aspect would crop.

| Sample slot | Cell | Frames | Dataset | Caption | Episode | Window | Sample id |
|---|---|---:|---|---|---|---:|---|
| `media/rollout_abc130k__s01__*.mp4` | 480x360 | 81 | abc130k | fold and stack the t-shirts | `fold_and_stack_the_t_shirts/episode_7de09792-96e5-4e59-b3b5-bbef6dd350f4 / top` | 891 | `8762897a20b3` |
| `media/rollout_abc130k__s02__*.mp4` | 480x360 | 81 | abc130k | place the food in the zip-top bag and seal the zipper | `place_the_food_in_the_zip_top_bag_and_seal_the_zipper/episode_70656bde-7b22-43f1-81be-248d03d930eb / top` | 891 | `bf1276a7302b` |
| `media/rollout_abc130k__s03__*.mp4` | 480x360 | 81 | abc130k | fold and stack the t-shirts | `fold_and_stack_the_t_shirts/episode_a570cd33-108f-41ab-8ee9-f96f43961f86 / top` | 162 | `dc3a72ee5776` |
| `media/rollout_deform360__s01__*.mp4` *(paper)* | 480x272 | 81 | deform360 | a human hand holding a robot gripper manipulating a rope | `001-rope/episode_3 / brics-odroid-015_cam0` | 0 | `1032eaed65c8` |
| `media/rollout_deform360__s02__*.mp4` | 480x272 | 81 | deform360 | a human hand holding a robot gripper manipulating a net cloth | `091-net-cloth/episode_4 / brics-odroid-024_cam0` | 81 | `bfc497ef1d38` |
| `media/rollout_deform360__s03__*.mp4` | 480x272 | 81 | deform360 | a human hand holding a robot gripper manipulating a cotton scarf cloth | `086-cotton-scarf-cloth/episode_0 / brics-odroid-015_cam1` | 81 | `e8ab8a7effa6` |
| `media/rollout_droid__s01__*.mp4` *(paper)* | 480x272 | 81 | droid_raw | fold, spread out, or clump object | `1.0.1/TRI/success/2023-10-25/Wed_Oct_25_17:11:21_2023 / ext1` | 162 | `08c1bc051bca` |
| `media/rollout_droid__s02__*.mp4` | 480x272 | 81 | droid_raw | Use the towel on the left to wipe the table | `1.0.1/REAL/success/2023-06-27/Tue_Jun_27_15:26:55_2023 / ext2` | 162 | `492a9788d932` |
| `media/rollout_droid__s03__*.mp4` | 480x272 | 81 | droid_raw | pnp-table-to-plate-onlycans | `1.0.1/RPL/success/2023-12-03/Sun_Dec__3_18:48:18_2023 / ext1` | 810 | `d1b96df38273` |
| `media/rollout_xvla_soft_fold__s01__*.mp4` *(paper)* | 480x360 | 81 | xvla_soft_fold | fold the cloth | `0706_17pm_stage_1_stage2new_new_cam_very_slow/episode_22 / cam_high` | 1539 | `2acae14159ae` |
| `media/rollout_xvla_soft_fold__s02__*.mp4` | 480x360 | 81 | xvla_soft_fold | fold the cloth | `0708_11am_stage_1_stage2new_new_cam_very_slow/episode_11 / cam_high` | 0 | `5bba6b196c19` |
| `media/rollout_xvla_soft_fold__s03__*.mp4` | 480x360 | 81 | xvla_soft_fold | fold the cloth | `0712_8pm_stage_1_stage2new_new_cam_very_slow/episode_14 / cam_high` | 810 | `c553f3bc75d4` |

## Figure 4 - running the interface backwards

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/control_generated_motion.mp4` | 640x418 | 137 | 16 | GEAR-YAM/wam_rollout/pipe_play, world-model generation conditioned on the transferred object flow | crop 62 rows off the bottom to drop the burned-in badge |
| `media/control_human_demo.mp4` | 640x418 | 107 | 15 | GEAR-YAM/wam_rollout/pipe_play, human demonstration, pipe_play | crop 62 rows off the bottom to drop the burned-in badge |
| `media/control_object_flow.mp4` | 640x418 | 107 | 15 | GEAR-YAM/wam_rollout/pipe_play, object flow derived from that demo | crop 62 rows off the bottom to drop the burned-in badge |
| `media/control_physical_execution.mp4` | 640x418 | 68 | 15 | GEAR-YAM/wam_rollout/pipe_play, physical YAM execution of the read-out actions | crop 62 rows off the bottom to drop the burned-in badge |

## Grounding and sampling

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/flow_dense_tracks.mp4` | 640x480 | 81 | 16 | AllTracker output before grounding, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | 611 tracks on a 34x19 image lattice, hue keyed to frame-0 position so the pool reads as unlabeled |
| `media/flow_embodiment_mask.mp4` | 640x480 | 81 | 16 | SAM3 embodiment masks (gripper_left u gripper_right), processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | per-frame mask, translucent fill plus outline |
| `media/flow_object_mask.mp4` | 640x480 | 81 | 16 | SAM3 primary_object mask, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | per-frame mask, translucent fill plus outline |
| `media/flow_sample_all.mp4` | 640x480 | 81 | 16 | all sampling mode over the same window, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | 256 tracks drawn by the figure's own selection rule |
| `media/flow_sample_embodiment.mp4` | 640x480 | 81 | 16 | embodiment sampling mode over the same window, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | 48 tracks drawn by the figure's own selection rule |
| `media/flow_sample_none.mp4` | 640x480 | 81 | 16 | conditioning dropout: the clip with no flow at all, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | none |
| `media/flow_sample_object.mp4` | 640x480 | 81 | 16 | object sampling mode over the same window, processed_v2 xvla_soft_fold episode_29/cam_high window 0, the window fig:action_flow_sampling is cut from | 48 tracks drawn by the figure's own selection rule |

## Method - offline training

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/method_train_flow_deform360.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, Deform360 | per-track trail over a 16-frame rolling history, then the same 4:3 crop |
| `media/method_train_flow_droid.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, DROID | per-track trail over a 16-frame rolling history, then the same 4:3 crop |
| `media/method_train_flow_xvla.mp4` | 640x480 | 81 | 16 | AllTracker tracks x grounded masks over the same window, XVLA-Soft-Fold | per-track trail over a 16-frame rolling history, then the same 4:3 crop |
| `media/method_train_video_deform360.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, Deform360 (fig:method_overview top lane) | centre-crop to 4:3, resize to 640x480 |
| `media/method_train_video_droid.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, DROID (fig:method_overview top lane) | centre-crop to 4:3, resize to 640x480 |
| `media/method_train_video_xvla.mp4` | 640x480 | 81 | 16 | processed_v2 intermediate window, XVLA-Soft-Fold (fig:method_overview top lane) | centre-crop to 4:3, resize to 640x480 |

## Method - online deployment

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/method_deploy_flow.mp4` | 640x480 | 277 | 16 | URDF forward kinematics of the policy's own states, projected through the calibrated camera (GEAR-YAM ep_000000), drawn first over the Isaac Lab replay of the same command | trails grow over a background that cross-fades from the Isaac replay to the held observation |
| `media/method_deploy_isaac.mp4` | 640x480 | 277 | 16 | Isaac Lab kinematic replay of the same policy rollout (real_world/isaaclab_replay.py, RTX raytraced, calibrated top D405) | native 640x480 render, no crop |
| `media/method_deploy_prediction.mp4` | 640x480 | 277 | 16 | closed-loop world-model generation, GEAR-YAM policy_in_wm/ep_000000 | none, native 640x480 |

## Policy evaluation - side-by-side replay

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/policy_robolab_failure_reference.mp4` | 856x442 | 277 | 16 | rated RoboLab episode, failure row, re-rendered without the conditioning overlay -- RoboLab reference pane | split at the midpoint, crop 20 rows off the top and 50 off the bottom to drop the burned-in caption and badge |
| `media/policy_robolab_failure_simulated.mp4` | 856x442 | 277 | 16 | rated RoboLab episode, failure row, re-rendered without the conditioning overlay -- our neural simulator pane | split at the midpoint, crop 20 rows off the top and 50 off the bottom to drop the burned-in caption and badge |
| `media/policy_robolab_success_reference.mp4` | 856x442 | 109 | 16 | rated RoboLab episode, success row, re-rendered without the conditioning overlay -- RoboLab reference pane | split at the midpoint, crop 20 rows off the top and 50 off the bottom to drop the burned-in caption and badge |
| `media/policy_robolab_success_simulated.mp4` | 856x442 | 109 | 16 | rated RoboLab episode, success row, re-rendered without the conditioning overlay -- our neural simulator pane | split at the midpoint, crop 20 rows off the top and 50 off the bottom to drop the burned-in caption and badge |

## Wrist-camera egomotion

| Slot | Size | Frames | fps | Source | Transform |
|---|---|---:|---:|---|---|
| `media/wrist_camera_egomotion.mp4` | 1296x284 | 17 | 3 | outputs/wan22_droid_wrist_a14b, one DROID wrist 17-frame window | three labelled panes side by side, native 848x480 each |

## Action-flow overlays

`teaser_flow_*` draws each track's trail up to the current frame, on that
frame, so the clip and its conditioning signal advance together. Tracks
come from the cohort's own AllTracker window and are classified by the
training reader's rule -- frame-0 position against the frame-0 grounded
masks, embodiment taking precedence over object. Colour is per-track and
constant in time (measured to be the convention the shipped stills use)
and encodes the class the captions name. Each class draws from a band
rather than one flat value -- embodiment H 0-34, object H 100-158, with saturation and value
ramped alongside -- keyed to the track's frame-0 position within its own
class extent, so neighbouring tracks take neighbouring tones and a
two-armed rollout separates into a red arm and an amber one. The bands are
centred on the flat key colours (embodiment `rgb(255, 96, 48)`, object `rgb(72, 224, 136)`),
which are what the page's legend swatches and the mask fills still use.
Density is one track per lattice cell
(40 px embodiment, 36 px object).

`method_train_flow_*` draws the same way but keeps only the last
16 frames of each trail: these three windows were chosen for
large motion, and at full history the accumulated trails cover the arm
and the cloth they are tracking by the middle of the window. The
deployment lane keeps full history on purpose -- its observation is
held, so the growing trail *is* the commanded trajectory. That lane's
clip opens on the Isaac Lab replay and dissolves to the observation over
frames 32-48: both are the same
calibrated top D405, so the projected tracks sit on the arms in either,
and the page's step 2 -> step 3 hand-off is shown rather than asserted.

One deviation from the paper. `fig:action_flow_sampling` colours
embodiment `#76B900` and object `#F09039`, i.e. the opposite way round
from every other figure; the page would then use green for two
different things a screen apart. The sampling clips are re-rendered in
the page-wide palette above and the page carries an explicit key.

| Slot | Tracks in window | Embodiment | Object | Drawn (emb / obj) |
|---|---:|---:|---:|---|
| `media/method_deploy_flow.mp4` | 365 | 365 | 0 | 98 / 0 |
| `media/method_train_flow_deform360.mp4` | 16004 | 1045 | 1051 | 29 / 37 |
| `media/method_train_flow_droid.mp4` | 16143 | 3347 | 0 | 71 / 0 |
| `media/method_train_flow_xvla.mp4` | 16094 | 2363 | 1238 | 58 / 31 |
| `media/teaser_flow_deform360.mp4` | 16150 | 555 | 1151 | 22 / 47 |
| `media/teaser_flow_droid.mp4` | 16374 | 1683 | 0 | 47 / 0 |
| `media/teaser_flow_egodex.mp4` | 16182 | 504 | 0 | 15 / 0 |
| `media/teaser_flow_xvla_yam.mp4` | 16384 | 1731 | 2111 | 44 / 41 |
