How it works
From “program the robot” to “train the robot”
An industrial arm has six joints and runs the same path a million times. A humanoid has thirty to fifty joints, two hands, a camera head, and a job that changes every shift. You cannot script that. You teach it.
The old way, and why it broke
Classical robotics is motion planning: a human engineer defines the task, a planner computes a joint trajectory, and a controller executes it. This works beautifully when the world is fixed. Welding a car body is a fixed world.
A humanoid walking into a parts cage, picking a tote whose contents shift, and placing it on a cart that a person left slightly askew is not a fixed world. Every variation is a new program. The engineering cost explodes.
So the field borrowed the trick that made large language models work: stop writing rules, collect a very large number of examples, and train one model to map what the robot sees and is told to what the robot should do next.
The result is called a policy. In 2026 the dominant form of policy is the vision-language-action model, or VLA. Covered on the Models page.

The four-layer stack
- Layer 1 — Pre-train on human videoMillions of hours of people cooking, assembling, carrying. No robot in the frame. The model learns what objects are, how scenes are laid out, roughly how hands move. Cheap and enormous, but it says nothing about this robot’s joints.
- Layer 2 — Pre-train in simulationA physics engine such as NVIDIA Isaac Sim or MuJoCo runs thousands of robots in parallel. Locomotion, balance, recovery from a shove: these are learned almost entirely in simulation, because falling over ten thousand times is free there and ruinous in a plant.
- Layer 3 — Fine-tune on teleoperationA person wears a VR headset or an exoskeleton and drives the actual robot through the actual task. Every second produces a matched pair: what the cameras saw, what the joints did. This is the expensive, slow, irreplaceable layer. It is where a factory’s own tasks get into the model.
- Layer 4 — Deploy and keep learningThe robot runs the task. When it fails, a remote operator takes over, and that intervention becomes new training data. The fleet gets better together. 1X Technologies calls this “Expert Mode”; most companies have a version of it.

The number that makes it real
5–50 episodes per operator-hour. That is the throughput of teleoperation today: a skilled person can produce somewhere between five and fifty usable demonstrations of a task in an hour. Figure’s Helix model was reportedly trained on roughly 500 hours of such data. NVIDIA’s answer was to take a few dozen human demonstrations and multiply them into 540,000 synthetic ones with a tool called DexMimicGen. The whole industry is a race to get more Layer 3 quality out of less Layer 3 time.
The caveat nobody puts in the press release
The sim-to-real gap. A policy that is perfect in simulation drops in performance the moment it meets real friction, real sensor noise, and real lighting. Contact-rich tasks, anything involving two hands and a deformable object, suffer most. The reliable fix is real teleoperation data layered on top of synthetic pre-training. There is no shortcut yet, which is why the teleoperation companies on the Data page exist.