Models
Vision-language-action models, the brain that got trained
A VLA takes in camera frames and a sentence, and outputs the next joint positions. One model, end to end. In 2026 nearly every humanoid company has one, and the open ones are good enough to fine-tune in a week.
The ones that matter
| Model | Who | Open? | What is notable |
|---|---|---|---|
| Helix | Figure AI | No | Controls the entire upper body of a humanoid, including head gaze, wrists, torso and individual fingers, at high rate. Reportedly trained on about 500 hours of multi-robot data. Runs the robots at BMW. |
| GR00T N1 → N1.5 → N1.6 → N1.7 | NVIDIA | Yes | The open foundation model for humanoids. N1 was a 2B-parameter model trained on real humanoid data plus synthetic data from Omniverse and Cosmos. N1.7 reached early access with commercial licensing in late 2026, aimed at production deployments, with AGIBOT, Humanoid, LG Electronics, NEURA Robotics and Noble Machines named as adopters. |
| GR00T N2 | NVIDIA | Previewed | Previewed at GTC and slated for availability by the end of 2026. Built on a world action model framework from the DreamZero research rather than a conventional VLA stack, and reported to succeed at new tasks in new environments more than twice as often as leading VLAs. Currently top of MolmoSpaces and RoboArena, two public leaderboards that score a single policy across many tasks and robot bodies rather than on one benchmark task. |
| Gemini Robotics → Gemini Robotics 2 | Google DeepMind | No | Brings a frontier language model into the loop. Gemini Robotics 2, announced July 2026, extends to whole-body intelligence. |
| π0 → π0.5 → π0.7 | Physical Intelligence | Partly | Flow-matching VLA family with open-world generalisation. π0.7 (April 2026) is a steerable generalist model reporting zero-shot cross-embodiment generalisation and emergent capabilities. |
| 1X World Model | 1X Technologies | No | A world model used to evaluate and train the NEO home humanoid, predicting what will happen before the robot acts. |
| Large Behavior Models | Boston Dynamics / TRI | No | The policy family behind the electric Atlas, trained on teleoperated demonstrations. |
| Xiaomi-Robotics-1 | Xiaomi | No | Trained on over 100,000 hours of real-world manipulation trajectories collected via UMI devices. The clearest demonstration so far that the text scaling curve may hold for motion. |

How a VLA is built, in one paragraph
Take a vision-language model that already understands pictures and sentences. Bolt on an action head, usually a diffusion or flow-matching network, that turns the model’s understanding into a short sequence of future joint targets, sixteen steps at a time in GR00T’s case. Pre-train on human video and simulation. Fine-tune on teleoperation. The “slow” language brain runs at a few hertz; the “fast” action head runs at tens to hundreds of hertz. Figure calls the two halves System 2 and System 1. Most of the field has copied the shape.

What is changing in 2026
World models as data engines
Instead of only recording what happened, generate what could happen. Video world models produce plausible futures of a scene, which become training data. NVIDIA’s Cosmos and a wave of 2026 papers push this line. The 1X World Model is the clearest commercial example.
Training on zero robot data
Several 2026 models claim a working policy trained with no robot data at all, only human video retargeted onto the robot. Sunday Robotics’ ACT-1 is one. If this holds up on factory tasks, the teleoperation bottleneck becomes a fine-tuning step rather than the main event.
Reinforcement fine-tuning in simulators
After imitation, let the policy practise in a world simulator with verified rewards. The policy gets better than its demonstrations, something pure imitation cannot do.
Scale, measured in hours
Xiaomi’s 2026 model reports more than 100,000 hours of real-world trajectories. Two years ago, 500 hours was a headline number. The scaling curve that worked for text is being tested on motion.