
The 4% problem: why most egocentric robot training data fails quality checks
Learn why most egocentric robot training data fails quality checks and how verified, deeply labeled human demonstrations can deliver greater training value than larger datasets.
7 min read time
Topics
Physical AI
Robotics
Robotics AI
AI Training Data
Humanoid Robotics
Verticals
Physical AI
Author(s)

Leela Krishna
Ask a robotics lab how much data it takes to teach a humanoid to fold a towel and you will hear a number in the millions of hours. But only a small fraction of a large egocentric video collection may be suitable for training. The 2026 Do as I Do study sampled 2,000 ten-second clips from a corpus already filtered for hand-object contact. Only 9% contained meaningful interaction, and just 83 clips, or 4%, survived the full 3D hand-object reconstruction process.

That low yield changes the economics of collecting robot training data. Teleoperation requires a skilled operator for every robot hour. Egocentric capture can collect demonstrations without operating a robot, but the resulting footage still has to pass quality checks before it can be used for training. The relevant measure is not simply how many hours a team collects, but how many verified hours are suitable for training and what each of those hours can teach a model.
Why egocentric robot training data has a quality problem
The 4% yield reported in the Do as I Do study is one result among several showing how much egocentric footage can be lost as researchers impose requirements for training use. The table below shows what four studies reported:
Table title: Egocentric video yield after training-data quality filters
Source | What was measured | Yield |
|---|---|---|
Do as I Do · 2606.19333 | 2,000 pre-filtered clips → clips passing full 3D hand-object reconstruction QC | 4% |
Ego4D · 2110.07058 | 3,670 h collected → hours annotated for hands and objects (2D only; no 3D hand pose in the corpus) | 5.3% |
ACE-Ego-0 · 2606.17200 | A 2026 pretraining pool after ego-view, manipulation and hand-confidence filters: 216.6 h kept from Ego4D, but 776.8 of 829 h kept from headset-tracked EgoDex | 5.9% vs 94% |
Being-H0 · 2507.15597 | 1,100 h of 3D-hand-pose training data; Ego4D and EPIC-KITCHENS contribute zero hours, excluded for “hand occlusion, out-of-frame, or dynamic camera viewpoints” | 0% |
ACE-Ego-0 shows how strongly the capture method affects yield. The same pipeline retained 5.9% of Ego4D but 94% of EgoDex, which uses headset-based hand tracking. What qualifies as usable data also affects yield. A filter that asks only whether a hand is visible can retain much more footage: MINT reports 59%. A filter that requires a reliable 3D hand pose, an in-frame interaction, a stable camera, and a verified task retains about one clip in twenty.
The number of hours in a dataset therefore says little about how many of those hours have been verified for training use. Yet most of the industry still prices data by total hours rather than verified hours.

Robot training data needs numerical and semantic quality checks
Every recording should pass two independent checks before it enters a training dataset. The numerical gate covers what most pipelines half-do: frame timing, sensor sync, hand-pose confidence with explicit truncation and occlusion flags, and SLAM-composed wrist tracks so hand motion is measured against the world rather than against a moving head. The semantic gate asks whether the task actually happened: did the object end where the instruction said, did contact occur when the label says it did, was this a success, a failure, or a failure with a recovery. Vision-language judges and human review answer together, and the verdict is stored on the episode as evidence, not a checkbox.
A verified episode is still just video with a joint trace. To be useful for mid-training a VLA or a world action model, it needs phase segmentation (reach, grasp, transport, place, release), object-centric flow tracks that transfer across embodiments, contact and outcome events with failure and recovery as first-class labels, and language at two grains, one instruction per episode and one caption per phase. The same labelled episode feeds both model families, which is the economic case for annotating deeply rather than collecting widely.

Why verified egocentric data can outperform larger datasets
Several recent studies have compared the training value of egocentric human video with larger quantities of robot data:
HumanEgo (2026): 30 minutes of human egocentric video per task reached 92.5% success versus about 51% for 30 minutes of robot teleoperation on the same tasks. Eight minutes of human video already beat 30 minutes of teleop. Swapping raw pixels for an entity-level hand-object representation moved success from 7.5% to 85%. The labels are the lever.
HumanScale (2026): at a matched 5,000 hours, egocentric human video beat real-robot data for pretraining, with 52.5% higher in-distribution and 90% higher out-of-distribution success on a real robot. A human hour holds about five times as many trajectories as a teleop hour.
EgoScale (2026): mid-training on 50 hours of human data and only 4 hours of robot data; one robot demonstration plus 100 aligned human demonstrations reached 0.88 on shirt folding.
Quality over Quantity (2026): keeping only the top half of demonstrations, ranked by their influence on the policy, raised real-robot success from 56.7% to 86.7%. Throwing away the worse half made the policy better.
Unconstrained capture yields relatively few verified hours, but the studies above show how much training value those hours can contain. Densely labelled human data can outperform larger quantities of teleoperation data, while curation can improve results further. The economics favor verifying and labelling the data that survives collection rather than simply increasing the number of hours collected. Andrew Ng’s line about fifty thoughtfully engineered examples was written for language models. It applies here with fewer zeros.
Some of the labs that demonstrated the benefits of larger datasets are now focusing more closely on data quality. One frontier embodied-model group that published a 278,000-hour scaling result has reported that pretraining data quality persists into fine-tuning, that task completion matters more than validation loss, and that its latest model learns a new short task from a single 3 second-to–12 second demonstration. When one good demonstration is enough, the supplier's job is to supply that one, plus the failures and recoveries that teach what “good” looks like.
How Humanoid IQ builds a curated robot training task library
This is the design principle behind Humanoid IQ, Centific’s platform for training humanoids, VLAs and world action models from a curated task library rather than a pile of hours. Every task is chosen and recorded to transfer along four axes: embodiment (two-finger grippers to 21- and 22-DoF hands, with human hands as the source), capture modality (egocentric, egocentric stereo, VR and full-body teleop, paired on the same task), vertical (twelve domains mapped to a task knowledge graph built on O*NET, GRASP and BEHAVIOR-1K), and outcome mix (roughly 70% success, 10% failure, 20% failure with recovery). Every accepted demonstration is replayed on a virtual robot before it ships, and the Data Marketplace exposes the library with free samples.
A team building a VLA or a WAM does not need 300,000 hours. It needs the right few hundred tasks, each with a verified success, a few instructive failures, and the recoveries that connect them, in one training-ready format.
Inspect Humanoid IQ training data yourself
Until now, the only way to know whether a dataset would survive this kind of scrutiny was to license it and find out. We have launched centific.com/physical-ai, where anyone can open real Humanoid IQ episodes, human and robot, and run the same numerical and semantic checks we run internally. Step through an episode frame by frame. See the hand-pose confidence and truncation flags, the SLAM-composed wrist track, the phase boundaries, the contact events, and the verdict on whether the task was completed. Then decide for yourself whether it is an hour you would train on.
Try it yourself. Open the inspector with free sample data at robot-iq.centific.com/inspect. To run the checks on your own recordings, book a demo.
Dataset size means little without knowing how many hours have been verified. Humanoid IQ makes that verification visible so teams can inspect the data and decide whether they would train on it.
Learn more about Centific’s Physical AI capabilities
Explore Physical AI at Centific, or see From raw robotics data to training-ready AI datasets and Training humanoids for actions that matter.
Are your ready to get
modular
AI solutions delivered?
Connect data, models, and people — in one enterprise-ready platform.
Latest Insights
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.

