When you order robot learning data, the format it arrives in decides how much work sits between delivery and your first training run. Two formats come up constantly: RLDS, from Google's research ecosystem, and LeRobotDataset, from Hugging Face's LeRobot project. Here is what each one is, how they differ, and what to put in your request so the data lands ready to use.
Why the format matters
Robot data is not just video. A single demonstration can include several camera streams, joint positions, gripper state, commanded actions, timestamps, a language instruction and success labels. All of that has to stay aligned frame by frame, grouped into episodes, and described consistently across thousands of recordings.
If a vendor sends a folder of MP4s and loose CSVs, your team ends up writing conversion scripts, guessing at synchronization and debugging off-by-one errors between observations and actions. A standard format removes most of that. It also lets you mix your data with public datasets and reuse existing training code.
RLDS: episodes of steps
RLDS (Reinforcement Learning Datasets) is an ecosystem from Google Research for recording, sharing and using sequential decision-making data. It is built on TensorFlow Datasets (TFDS), and data is loaded as tf.data pipelines.
The structure is simple:
- A dataset is a collection of episodes. Each episode has its own metadata (for example an episode ID or environment configuration) and a sequence of steps.
- Each step has two required boolean fields, is_first and is_last, plus optional fields such as observation, action, reward, discount and is_terminal. Custom fields can be added.
- All steps in a dataset must share the same fields, which is what makes datasets interchangeable across tools.
One detail worth knowing: in RLDS, a step pairs an observation with the action taken in response to it. On the final step (is_last), action and reward are not meaningful. If a vendor builds RLDS episodes by hand, check that this alignment is correct, because a one-step shift silently corrupts training.
RLDS matters in practice because of Open X-Embodiment, the large cross-robot collection that pooled data from many labs and spans 22 robot embodiments. Its datasets use the RLDS episode format, so if your training code was built around that collection or models trained on it, RLDS is the natural fit.
LeRobotDataset: Parquet plus video
LeRobotDataset is the format used by Hugging Face's LeRobot library, and it is built around the Hugging Face Hub and PyTorch. Its current version, v3.0, splits data into three parts:
- Tabular data in Parquet: low-dimensional, high-frequency signals such as robot state, actions and timestamps.
- Video in MP4: camera frames encoded per camera, rather than stored as individual images.
- Metadata in JSON and Parquet: a meta/info.json file with the schema (feature names, data types, shapes), frame rate and path templates; stats.json with normalization statistics; a tasks file mapping task descriptions to IDs; and per-episode records with lengths and offsets.
Features follow a naming convention, for example observation.state, action, observation.images.<camera_name> and timestamp. Loading a frame returns a dictionary of PyTorch tensors, and the library can return windows of past frames using time offsets, which is convenient for policies that look at short histories.
The main change from v2.x is storage layout. Earlier versions stored roughly one file per episode. v3.0 packs many episodes into larger Parquet and MP4 files and uses metadata offsets to find episode boundaries. That reduces file count, speeds up loading at scale and supports streaming directly from the Hub. A converter from v2.1 exists, but if you are starting fresh, ask for v3.0 unless your code is pinned to an older version.
Which one should you ask for?
The honest answer is: whichever your training stack already reads. Converting between them is possible, but it is easier to get it right once at the source.
| RLDS | LeRobotDataset | |
|---|---|---|
| Ecosystem | TensorFlow / TFDS | PyTorch / Hugging Face Hub |
| Images | Stored inside step records | MP4 video per camera |
| Low-dim signals | Step fields | Parquet tables |
| Good fit when | You build on Open X-Embodiment style pipelines | You train with LeRobot or share via the Hub |
For data that is not robot teleoperation, such as human egocentric video, there are no robot actions to record. In that case, a plain delivery of MP4 with JSON metadata and a CSV manifest is often the best starting point, and you can map it into either format later. Many teams ask for both: raw files for archive and inspection, plus a packaged version for training. See what egocentric video is and why robotics teams use it for context.
What to put in your request
Whatever format you choose, these details should be written down before collection starts. They belong in your data collection spec.
- Format and version. "LeRobotDataset v3.0" or "RLDS via TFDS", not just "robot format".
- Feature schema. Exact names, data types and shapes for each observation and action field, including units (radians or degrees, meters or millimeters) and coordinate frames.
- Camera list. Camera names, resolution, frame rate and mounting position for each stream.
- Timing and synchronization. Target control frequency, how timestamps are generated and how much drift between streams is acceptable.
- Episode definition. What counts as the start and end of an episode, how failed or partial attempts are handled, and whether they are flagged or removed.
- Language and labels. Task instructions per episode, success flags, and any step-level annotations.
- Normalization stats. Whether the vendor should compute them, and over which split.
- Sample first. A small sample that loads cleanly in your own code before the full batch is produced.
The last point is the most important. A sample that loads, plays back correctly and trains for a few steps will catch nearly every schema mistake early. That is the core of a good data pilot.
Common mistakes to avoid
- Action and observation misalignment. Check that each action really follows the observation it is paired with.
- Silent unit changes. Mixed units across sessions break training without any obvious error.
- Inconsistent camera names. "wrist", "wrist_cam" and "cam_wrist" in different episodes turns one feature into three.
- Missing metadata. Without contributor, environment and device fields, you cannot audit coverage or debug a bad batch.
Masana delivers MP4 and WAV with JSON metadata and a CSV manifest by default, and packages data as RLDS or LeRobot on request, with a sample checked against your loader before scaling.
Want data that loads in your training code on day one? See how we scope, pilot and deliver collections in the format you need.
See how it works