Map dataset cameras onto OpenPI slots and reuse v3 conversions - #1
Merged
Merged
Conversation
--delta_joint_actions still leaves only the last dimension absolute. --absolute_action_dims replaces that set for robots with more than one gripper. Delta runs store norm stats of the chunk deltas the model sees, resume reuses checkpoint stats, and constant quantile bins are widened. Multi-hour episodes get exact per-frame timestamps, and the loader tolerance is half a frame so float32 rounding does not abort training while a dropped frame still does.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Camera selection. Pi0 / Pi0.5 have three fixed image slots. Keys are now matched by name (
front/base/high/exterior→base_0_rgb,wrist→left_wrist_0_rgb,right+wrist→right_wrist_0_rgb), and leftover keys fill empty slots. Unfilled slots stay zero images withimage_mask=false.--cameras,--drop_cameras(key or substring), and--camera_mapcontrol this. A dropped camera is removed from the dataset metadata, so LeRobot never decodes it. Asking for more than three cameras names any key that has no base/wrist role.Conversion is reused, and the source is left alone. v3 → v2 lands in
--convert_dirwith a completion stamp; a later run over the same videos reuses the tree instead of re-running ffmpeg over every episode. A camera subset becomes a sibling view of relative links plus a rewrittenmeta/info.json. A read-only dataset mount is staged into the cache rather than edited in place, so-v dataset:/data/input:roworks.Episode clips start at time 0.
-ssbefore-iwith-c copyleft the first PTS about one frame late, which LeRobot rejects (tolerance_s=1e-4). Asettsbitstream filter moves the copied timestamps to zero.Long episodes stay loadable. Parquet timestamps that no longer step by
1/fpsinside 1e-4 s (float32 on a multi-hour episode) are rewritten toframe / fpsin float64. The pinned LeRobot then casts those values back to float32, so the loader tolerance becomes half a frame. A dropped frame, whose error is1/fps, still fails. An explicittolerance_sabove 1e-4 is left unchanged.Task text and norm stats. Tasks are read from
tasks.parquetandtasks.jsonl. Norm stats read every state/action row by default (--norm_stats_max_frames 0); a positive cap samples a seeded subset of files. Results are folded in submission order, becauseRunningStatsis order-sensitive (running mean, and a quantile histogram anchored on the first batch). When--delta_joint_actionsis on, action stats are the same chunk deltas the training transform produces.--resumereusesnorm_stats.jsonfrom the newest checkpoint instead of recomputing. Quantile bounds withq99 - q01 < 1e-6are widened around the mean; mean and std are not changed.Checkpoints and logs.
--keep_perioddefaults to--save_interval, so every saved checkpoint survives; previously the manager'smax_to_keep=1kept only the last one and pruned the rest.--resumecontinues from the newest checkpoint. Per-step metrics are written line-buffered, soStep N: loss=...reachesdocker logsand log files instead of sitting in an 8 KiB stdout buffer for hours.Portability. Links between the cache and its camera view are relative, and whole-video episodes are hardlinked (or copied across mounts) instead of pointing back at the dataset mount, so the trees stay readable outside the container that built them.
Compatibility
Existing commands keep their meaning. Omitting
--delta_joint_actionsstill trains absolute actions. Passing only--delta_joint_actionsstill keeps the last dimension absolute.--absolute_action_dimsis optional and errors if used alone; with the delta flag it replaces that last-dimension set (indices oraction.names). Checkpoint assets already written are what--resumeloads. Floating image tags move only when this merges tomain. Publishedv*tags are not retagged.Flags added
--dataset_dir,--output_dir,--run_name,--exp_name,--convert_dir,--cameras,--drop_cameras,--camera_map,--delta_joint_actions,--absolute_action_dims,--keep_period,--resumeTest plan
pyteston the Pi0.5 image's Python 3.11 — new tests for the delta mask, chunk-delta stats, quantile widening, resume stats, the unmapped-camera error, and timestamp repairtestplus both image builds (PR builds do not push)camera_low,--absolute_action_dims right_gripper,left_gripper, 50 steps, batch 8. Step 0 loss 0.153, grad norm 4.59, checkpoint at step 49