How Humanoids Learn
How Humanoids Learn
A field guide · 16 parts · 112 concepts · snapshot as of September 30, 2026

How humanoids learn

A modern humanoid is a stack of learned controllers running at different speeds: a reasoning model a few times a second, a visuomotor policy hundreds of times a second, a balance controller up to a thousand. This guide builds that stack from first principles, one concept at a time, then follows the data, tools, companies and money that turn it into products. Almost every section has a working model you can poke at.

Fig 0.1 · the stack at a glance
Loading figure…
One second in the life of a humanoid. Each layer ticks at its own rate; slower layers run bigger models and make broader decisions. Tap a layer to read about it.

Four lenses on every concept

Each of the 112 concepts is written the same way, so you can read for the idea, the motivation, the trade-offs or the history, and skip what you already know.

From first principles

What it is and how it works, built up from things defined earlier. Equations appear only when they carry the argument, with a worked example.

Why it’s needed

The specific problem it solves, and what breaks without it.

Versus the alternatives

The competing approaches, with what each wins and loses. The concept itself is marked so you can compare like with like.

How we got here

A dated timeline of the papers and releases that led to it, each linked to its source.

Two tracks

Principles explains how a humanoid learns, from feedback loops to foundation models; it changes slowly. Practice is a dated snapshot of how the field actually gathers data and trains robots: capture rigs, teleoperation, NVIDIA’s stack, the labs, the market. Practice concepts lean on Principles through chips like B7, so you can start in either track and jump back when a term is new. Short on time? Part P sums up the Practice track in five concepts, and the directory maps who does what.

Principles

A–J · how it works

63 concepts · read in order
Part A7 concepts

Foundations

The vocabulary every later part stacks on: what a policy is, why a humanoid can’t simply be pushed where you want it, how motors are really commanded, and the neural-network words used throughout.

A1A2A3A4A5A6A7
Part B9 concepts

Reinforcement learning for bodies

How walking, balance and whole-body skills are learned by trial and error in simulation, and how those skills survive the jump to real hardware.

B1B2B3B4B5B6B7B8B9
Part C6 concepts

Motion data and motion tracking

Where human movement comes from, how it is mapped onto a robot body, and how a policy learns to perform it on a real, falling-prone machine. This part ends at the learned whole-body controller that everything above it depends on.

C1C2C3C4C5C6
Part D13 concepts

Manipulation policies and VLAs

How robots learn to use their hands from human demonstrations: why naive copying fails, how generative models fixed it, and how those models were grafted onto web-scale vision-language models to become VLAs.

D1D2D3D4D5D6D7D8D9D10D11D12D13
Part E6 concepts

RL post-training for manipulation

Imitation gets a robot to “usually works”. Deployment needs “almost always works”. This part covers how robots improve from their own attempts, successes and failures included.

E1E2E3E4E5E6
Part F7 concepts

Video, world models and latent actions

The internet has billions of hours of video and almost no robot actions. This part covers the ways the field turns video into robot training signal: inventing actions for it, generating new video, predicting in latent space, and training policies that imagine the future.

F1F2F3F4F5F6F7
Part G8 concepts

Data

Every method in this guide is ultimately limited by data: how much, how diverse, and how close to the robot’s own body. This part covers where robot data comes from, what it costs, and what is known about how much of it you need.

G1G2G3G4G5G6G7G8
Part H4 concepts

Evaluation and infrastructure

How do you know a policy got better? This part covers the statistics of robot evaluation, the benchmarks labs compare on, the formats that carry robot data, and the compute and latency budgets every design has to fit.

H1H2H3H4
Part I2 concepts

Hardware co-design

Learning doesn’t happen in a vacuum. The actuators decide what motions are possible and how hard the robot is to simulate; the sensors decide what a policy can perceive. This part covers both choices and how they interact with learning.

I1I2
Part J1 concept

How it all connects

One household task, every layer of the stack: which concept does the work at each moment, at what rate, and which data trained it.

J1
Practice

K–P · how it’s done now

49 concepts · as of September 30, 2026
Part K9 concepts

Learning from people: egocentric capture

How first-person video of people becomes robot training data: the capture rigs, the pipeline from pixels to actions, the training recipes that make human data transfer, and what the evidence shows.

K1K2K3K4K5K6K7K8K9
Part L8 concepts

Teleoperation: the gold-standard data

How operators drive robots to record the robot’s own actions: interfaces, latency, the operations behind a data factory, the training recipe step by step, offline evaluation, and the loop that turns deployment into improvement.

L1L2L3L4L5L6L7L8
Part M12 concepts

NVIDIA’s robotics stack

NVIDIA doesn’t build robots; it sells the three computers every robot company needs. This part takes its stack apart layer by layer: physics, simulators, assets, data generation, world models, GR00T, whole-body control and Jetson Thor.

M1M2M3M4M5M6M7M8M9M10M11M12
Part N9 concepts

The labs and their bets

Every lab building a robot brain now agrees on the architecture and disagrees on the data. Who is betting on what, how each pipeline works end to end, and how strong the evidence is.

N1N2N3N4N5N6N7N8N9
Part O6 concepts

The market for robot data

Several hundred companies and state-backed centres now sell training inputs to robot makers. Who sells what, what it costs, where the money went in 2025–26, what China builds as state policy, who owns the data, and which segments are crowded or empty.

O1O2O3O4O5O6
Part P5 concepts

State of play, September 2026

The Practice track pulled together: how the four data sources compare, the research that moved the field, the recipe the leading labs converged on and the evidence for each bet, the open problems, and a checklist for builders and investors.

P1P2P3P4P5
Directory133 entries

Who’s who in humanoid learning

Robot makers, model labs, data and teleoperation vendors, simulation and tooling companies, and China’s state training grounds. Filter by layer, region and data strategy; every entry links to the concepts that explain its approach.

Open the directory →

A note on the numbers

Figures quoted from papers and company posts are linked in each timeline. Company results (Figure, 1X, Generalist, Physical Intelligence, NVIDIA and others) are mostly self-reported and not peer-reviewed; Part H explains how to read them. Practice-track claims carry an evidence tag: paper a paper or preprint with methods, docs open code, data or technical documentation, claim a company statement nobody has independently checked, reported a press report of a deal that has not closed, and unverified where the sources disagree or say nothing. Interactive figures marked Toy model use small illustrative models to show a mechanism, not to reproduce a published number.