Skip to content

Map dataset cameras onto OpenPI slots and reuse v3 conversions - #1

Merged
joaner merged 7 commits into
mainfrom
fix/camera-slots-and-v3-cache
Sep 23, 2026
Merged

joaner merged 7 commits into
mainfrom
fix/camera-slots-and-v3-cache

Conversation

@joaner

@joaner joaner commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

What this changes

Camera selection. Pi0 / Pi0.5 have three fixed image slots. Keys are now matched by name (front / base / high / exterior → base_0_rgb, wrist → left_wrist_0_rgb, right + wrist → right_wrist_0_rgb), and leftover keys fill empty slots. Unfilled slots stay zero images with image_mask=false. --cameras, --drop_cameras (key or substring), and --camera_map control this. A dropped camera is removed from the dataset metadata, so LeRobot never decodes it. Asking for more than three cameras names any key that has no base/wrist role.

Conversion is reused, and the source is left alone. v3 → v2 lands in --convert_dir with a completion stamp; a later run over the same videos reuses the tree instead of re-running ffmpeg over every episode. A camera subset becomes a sibling view of relative links plus a rewritten meta/info.json. A read-only dataset mount is staged into the cache rather than edited in place, so -v dataset:/data/input:ro works.

Episode clips start at time 0. -ss before -i with -c copy left the first PTS about one frame late, which LeRobot rejects (tolerance_s=1e-4). A setts bitstream filter moves the copied timestamps to zero.

Long episodes stay loadable. Parquet timestamps that no longer step by 1/fps inside 1e-4 s (float32 on a multi-hour episode) are rewritten to frame / fps in float64. The pinned LeRobot then casts those values back to float32, so the loader tolerance becomes half a frame. A dropped frame, whose error is 1/fps, still fails. An explicit tolerance_s above 1e-4 is left unchanged.

Task text and norm stats. Tasks are read from tasks.parquet and tasks.jsonl. Norm stats read every state/action row by default (--norm_stats_max_frames 0); a positive cap samples a seeded subset of files. Results are folded in submission order, because RunningStats is order-sensitive (running mean, and a quantile histogram anchored on the first batch). When --delta_joint_actions is on, action stats are the same chunk deltas the training transform produces. --resume reuses norm_stats.json from the newest checkpoint instead of recomputing. Quantile bounds with q99 - q01 < 1e-6 are widened around the mean; mean and std are not changed.

Checkpoints and logs. --keep_period defaults to --save_interval, so every saved checkpoint survives; previously the manager's max_to_keep=1 kept only the last one and pruned the rest. --resume continues from the newest checkpoint. Per-step metrics are written line-buffered, so Step N: loss=... reaches docker logs and log files instead of sitting in an 8 KiB stdout buffer for hours.

Portability. Links between the cache and its camera view are relative, and whole-video episodes are hardlinked (or copied across mounts) instead of pointing back at the dataset mount, so the trees stay readable outside the container that built them.

Compatibility

Existing commands keep their meaning. Omitting --delta_joint_actions still trains absolute actions. Passing only --delta_joint_actions still keeps the last dimension absolute. --absolute_action_dims is optional and errors if used alone; with the delta flag it replaces that last-dimension set (indices or action.names). Checkpoint assets already written are what --resume loads. Floating image tags move only when this merges to main. Published v* tags are not retagged.

Flags added

--dataset_dir, --output_dir, --run_name, --exp_name, --convert_dir, --cameras, --drop_cameras, --camera_map, --delta_joint_actions, --absolute_action_dims, --keep_period, --resume

Test plan

  • pytest on the Pi0.5 image's Python 3.11 — new tests for the delta mask, chunk-delta stats, quantile widening, resume stats, the unmapped-camera error, and timestamp repair
  • CI on this push: test plus both image builds (PR builds do not push)
  • Pi0.5 LoRA smoke on WipeTable (DualArxR5a, 14-D, four cameras): drop camera_low, --absolute_action_dims right_gripper,left_gripper, 50 steps, batch 8. Step 0 loss 0.153, grad norm 4.59, checkpoint at step 49
  • Earlier Pi0.5 LoRA on SO-101: 3 cameras then without front, 20000 steps, checkpoints kept

--delta_joint_actions still leaves only the last dimension absolute. --absolute_action_dims replaces that set for robots with more than one gripper. Delta runs store norm stats of the chunk deltas the model sees, resume reuses checkpoint stats, and constant quantile bins are widened. Multi-hour episodes get exact per-frame timestamps, and the loader tolerance is half a frame so float32 rounding does not abort training while a dropped frame still does.
@joaner
joaner merged commit 37f27b5 into main Sep 23, 2026
3 checks passed
@joaner
joaner deleted the fix/camera-slots-and-v3-cache branch September 23, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant