HumanoidTraining

Data

Where the training data comes from

Three sources, and in 2026 the answer is no longer “which one”. It is a stack: teleoperation for fidelity, simulation for scale, human video for scene understanding.

The three data sources compared
SourceWhat it isStrengthWeaknessBest used for
TeleoperationA person controls the real robot through a VR headset, leader arm, or full-body exoskeletonHighest fidelity; zero embodiment gapSlow: 5–50 episodes per operator-hour; needs the physical robotFine-tuning; contact-rich, two-handed tasks
SimulationPhysics engine (Isaac Sim, MuJoCo, Isaac Lab) runs thousands of robots at onceMassive scale at near-zero cost; safe failureSim-to-real gap in contact dynamics and visionPre-training; locomotion; edge cases and safety
Human videoEgocentric or third-person footage of people doing tasks; no robot presentEnormous volume; natural task varietyNo robot actions attached; must be retargetedPre-training; scene and object priors
Close-up of a hand in a thin tactile data glove gripping a metal component
Tactile gloves let the operator feel contact through the robot.
First-person view of hands sorting mechanical parts on a workbench
Egocentric human video: enormous volume, no robot actions attached.

Teleoperation, the irreplaceable layer

Teleoperation in 2026 sits at the centre of every serious humanoid program. It is no longer just a way to drive a robot. It is a structured data pipeline, and the patents say so: Boston Dynamics and Baker Hughes both describe XR teleoperation explicitly as a collection system for training autonomous manipulation policies.

The hardware has moved fast. Full-body systems now track 72 degrees of freedom at sub-millimetre precision. Tactile gloves let the operator feel contact and resistance through the robot’s hands, which matters because a model trained on demonstrations where the person could not feel the part learns to crush it.

A service industry has grown around it: teleoperation data labelling platforms, operator staffing, XR teleoperation software for arms and humanoids. Sanctuary AI reports learning new tasks in under 24 hours by combining teleoperation with Isaac Lab simulation.

The frontier research in 2026 is “robot-free” demonstration: wearable interfaces that let a person demonstrate a task with their own body and have it retargeted onto the humanoid afterwards, without the robot present. HumanoidUMI describes exactly this. If it works at scale, the operator-hour bottleneck loosens.

Simulation, the multiplier

NVIDIA’s GR00T N1 report is the clearest public example of how simulation multiplies scarce human data. Engineers collected a few dozen source demonstrations by teleoperation using a Leap Motion hand tracker. DexMimicGen then generated 10,000 new demonstrations for each source-and-target receptacle pair across 54 combinations, producing 540,000 demonstrations for pre-training. On top of that, “neural trajectories” were produced by video generation models, turning generated video into training data.

This is the pattern across the industry: human data is the seed, simulation is the fertiliser. It is also why the GPU bill for training a humanoid looks like the GPU bill for training a language model.

Open datasets

Cross-embodiment datasets such as Open X-Embodiment and AgiBot World let a model pre-train on many different robots before it ever sees yours. The remaining step, always, is a custom teleoperation cell that anchors the policy to the target humanoid’s own kinematics. Humanoid and whole-body data, captured from motion-capture suits, exoskeletons and full-body teleoperation, is the fastest-growing data category in 2026.

Xiaomi’s 2026 model is the current high-water mark for scale: over 100,000 hours of real-world manipulation trajectories, collected with UMI devices rather than conventional teleoperation rigs.