Robot dataset acceptance: what to verify before scaling collection
An episode count is not a training-ready deliverable. Accept the task contract, robot configuration, multimodal timing, outcome labels, coverage and dataset revision before paying to scale collection.

The short answer: do not buy robot demonstration data by episode count alone. First accept whether every episode can be interpreted, replayed, filtered, trained and tied to an independent evaluation. China's Ministry of Industry and Information Technology has placed Technical Requirements for Humanoid Robot Data Collection in its 2026 industry-standard development plan. That is a standards project, not a published or effective standard, but it is a useful signal to turn data collection into a controlled engineering deliverable now.
What the Chinese standards project does—and does not—mean
The official MIIT third batch of 2026 industry-standard projects lists project 2026-0576T-SJ, Technical Requirements for Humanoid Robot Data Collection. It is a new 12-month drafting project under the ministry's humanoid-robot and embodied-intelligence standardisation committee. The plan does not publish acceptance clauses. Procurement and review documents should therefore say “standards project established,” not “standard implemented,” “compliant” or “certified.”
The practical signal travels beyond humanoids and beyond China: robot data is becoming a product with a declared object, process, quality boundary and configuration history. Teams can prepare without guessing future clauses by making today's datasets traceable and replayable.
Six acceptance layers: volume is only the outer shell
| Layer | Acceptance question | Minimum evidence | Common failure |
|---|---|---|---|
| Task and outcome | What starts, succeeds, fails or aborts the task? | Task ID, instruction, initial state, terminal conditions and outcome label | Reviewers assign different results to the same trajectory |
| Embodiment and configuration | Which robot and controlled configuration produced it? | Robot, tool, sensors, mounting, calibration, controller and software revisions | Camera or end effector changes while old frames and normalisation remain |
| Modalities and semantics | What do image, state, action and force fields mean? | Names, types, shapes, units, frames, sample rates and missing-value rules | Matching dimensions hide different units or action references |
| Timing and episode integrity | Do the streams describe the same physical event? | Clock source, timestamps, alignment, drops, latency and start/stop reasons | Video and action are offset although playback looks complete |
| Quality and coverage | Does collection cover reachable deployment states? | Scene, object, operator, disturbance and outcome distributions; review and exclusion rules | Many episodes repeat one camera pose, object and clean success path |
| Splits and revision | Can training results be reproduced and compared? | Train/validation/test rules, revision log, exclusions and dataset–checkpoint binding | Neighbouring trajectories or the same scene leak into test |
Readable is not yet acceptable
The official LeRobotDataset API is episode-aware and carries feature schemas, tasks, frame rate, episode metadata and statistics. Its dataset tools let a reviewer inspect camera streams, robot states and actions on a shared timeline. Those facilities are a strong format and review foundation. They do not, by themselves, prove that the task definition is correct, the timing error fits the application, or the collected distribution represents the intended workplace.
Quality does not mean deleting every failed attempt. LeRobot's human-in-the-loop workflow records interventions, corrections and recovery motions alongside autonomous segments to address states reached when policy errors compound. Research on data quality in imitation learning frames quality around distribution shift and shows that state diversity is not automatically beneficial in every setting. Failures and interventions need explicit classes and intended uses: neither silently mix them into demonstrations nor discard them indiscriminately.
A seven-step gate from pilot capture to released dataset
- Write the task contract. Freeze objects, tools, initial state, permitted actions, observable success, failure classes and abort conditions.
- Baseline one embodiment. Record tool, cameras, calibration, control rate and software for one robot. Never merge a changed configuration silently.
- Capture a small pilot. Prove integrity, alignment, replay, outcome review and exclusion before adding operators or stations.
- Use stratified review. Sample across operator, scene, task, outcome, configuration and disturbance—not only the newest or most attractive videos.
- Keep raw lineage. Every conversion, crop, annotation edit and exclusion should resolve back to its source episode and reason.
- Design the split before scaling. Hold out the station, object, collection batch or operator that represents the actual generalisation question. Do not divide one continuous session randomly across train and test.
- Release through independent rollouts. Bind dataset revision, training configuration, checkpoint, evaluation initial state and episode outcomes. Dataset QA passing is not policy deployment approval.
Coverage should answer a deployment question
The DROID project studies diverse real-world manipulation through distributed collection across scenes and tasks. It also released improved camera calibrations for part of the dataset in 2025. That update is a useful reminder: a dataset is not an immutable delivery folder. Calibration, annotation, exclusions and derived formats need explicit revision relationships. The LeRobot LIBERO documentation likewise asks result reports to pin the dataset revision and keep evaluation initial conditions controlled when comparing policies. Simulation specifics do not replace real-robot acceptance, but same revision, same conditions and reproducible comparison are equally valuable on hardware.
Boundary: this is a data-engineering preparation and delivery framework, not an interpretation of unreleased standard clauses. It makes no compliance, safety or performance claim for any dataset, policy or robot. Where recordings include people, location, speech, confidential process information or other identifiable data, the responsible parties must perform the separate contractual, legal, data-governance and security review applicable to their project.
Frequently asked questions
Does project 2026-0576T-SJ mean the Chinese standard is already effective?
No. It means Technical Requirements for Humanoid Robot Data Collection is in an industry-standard development plan. Its eventual text, designation, publication and effective status must be confirmed from later official notices.
How many robot demonstrations are enough?
There is no universal count independent of the task, embodiment, disturbances and target policy. Use a small pilot to prove schema, timing, labels and evaluation, then expand according to uncovered states and real-robot evaluation results.
Why not delete every failed trajectory?
Uncontrolled failed actions can contaminate demonstrations, but labelled deviations, interventions and recovery segments can define boundaries and teach recovery. Class and intended use matter more than a blanket keep-or-delete rule.
If LeRobot can load the dataset, is it training-ready?
No. Loading proves basic format compatibility. Acceptance still needs task semantics, units and frames, multimodal alignment, outcome labels, configuration lineage, coverage, split integrity and independent evaluation.
Sources
These primary sources support the material facts and engineering boundaries discussed above.
- 工业和信息化部 — 2026 年第三批行业标准制修订计划
- Hugging Face LeRobot — Dataset API and metadata
- Hugging Face LeRobot — Using Dataset Tools
- Hugging Face LeRobot — Human-in-the-Loop Data Collection
- Hugging Face LeRobot — LIBERO evaluation and dataset revisions
- DROID — A Large-Scale In-the-Wild Robot Manipulation Dataset
- Data Quality in Imitation Learning
Evaluating robot control, bimanual manipulation or a mobile platform?
Talk to Matrix Dimension →