HumanoidTraining

Models

Vision-language-action models, the brain that got trained

A VLA takes in camera frames and a sentence, and outputs the next joint positions. One model, end to end. In 2026 nearly every humanoid company has one, and the open ones are good enough to fine-tune in a week.

The ones that matter

Principal humanoid policy models, October 2026
ModelWhoOpen?What is notable
HelixFigure AINoControls the entire upper body of a humanoid, including head gaze, wrists, torso and individual fingers, at high rate. Reportedly trained on about 500 hours of multi-robot data. Runs the robots at BMW.
GR00T N1 → N1.5 → N1.6 → N1.7NVIDIAYesThe open foundation model for humanoids. N1 was a 2B-parameter model trained on real humanoid data plus synthetic data from Omniverse and Cosmos. N1.7 reached early access with commercial licensing in late 2026, aimed at production deployments, with AGIBOT, Humanoid, LG Electronics, NEURA Robotics and Noble Machines named as adopters.
GR00T N2NVIDIAPreviewedPreviewed at GTC and slated for availability by the end of 2026. Built on a world action model framework from the DreamZero research rather than a conventional VLA stack, and reported to succeed at new tasks in new environments more than twice as often as leading VLAs. Currently top of MolmoSpaces and RoboArena, two public leaderboards that score a single policy across many tasks and robot bodies rather than on one benchmark task.
Gemini Robotics → Gemini Robotics 2Google DeepMindNoBrings a frontier language model into the loop. Gemini Robotics 2, announced July 2026, extends to whole-body intelligence.
π0 → π0.5 → π0.7Physical IntelligencePartlyFlow-matching VLA family with open-world generalisation. π0.7 (April 2026) is a steerable generalist model reporting zero-shot cross-embodiment generalisation and emergent capabilities.
1X World Model1X TechnologiesNoA world model used to evaluate and train the NEO home humanoid, predicting what will happen before the robot acts.
Large Behavior ModelsBoston Dynamics / TRINoThe policy family behind the electric Atlas, trained on teleoperated demonstrations.
Xiaomi-Robotics-1XiaomiNoTrained on over 100,000 hours of real-world manipulation trajectories collected via UMI devices. The clearest demonstration so far that the text scaling curve may hold for motion.
A humanoid robot facing a workbench with faint graphic overlays of a camera frame and a motion trace
One model, end to end: camera frames and a sentence in, the next joint positions out.

How a VLA is built, in one paragraph

Take a vision-language model that already understands pictures and sentences. Bolt on an action head, usually a diffusion or flow-matching network, that turns the model’s understanding into a short sequence of future joint targets, sixteen steps at a time in GR00T’s case. Pre-train on human video and simulation. Fine-tune on teleoperation. The “slow” language brain runs at a few hertz; the “fast” action head runs at tens to hundreds of hertz. Figure calls the two halves System 2 and System 1. Most of the field has copied the shape.

A humanoid robot standing still while three faint translucent copies of it rehearse different reaches
A world model predicts what would happen before the robot moves, so rehearsal costs nothing.

What is changing in 2026

World models as data engines

Instead of only recording what happened, generate what could happen. Video world models produce plausible futures of a scene, which become training data. NVIDIA’s Cosmos and a wave of 2026 papers push this line. The 1X World Model is the clearest commercial example.

Training on zero robot data

Several 2026 models claim a working policy trained with no robot data at all, only human video retargeted onto the robot. Sunday Robotics’ ACT-1 is one. If this holds up on factory tasks, the teleoperation bottleneck becomes a fine-tuning step rather than the main event.

Reinforcement fine-tuning in simulators

After imitation, let the policy practise in a world simulator with verified rewards. The policy gets better than its demonstrations, something pure imitation cannot do.

Scale, measured in hours

Xiaomi’s 2026 model reports more than 100,000 hours of real-world trajectories. Two years ago, 500 hours was a headline number. The scaling curve that worked for text is being tested on motion.