Foundations
The vocabulary every later part stacks on: what a policy is, why a humanoid can’t simply be pushed where you want it, how motors are really commanded, and the neural-network words used throughout.
Robot-learning papers assume a shared vocabulary that no single paper defines. This part builds it from zero. If you already know what a PD loop, a floating base and a transformer are, skim the lede of each concept and move on to B1.
State, observation, action, policy
A robot never sees the world directly. It sees its sensors, and a policy is the rule that turns what it sees into what it does, re-applied many times a second.
MechanismFrom first principles
Start with the smallest possible robot: a cart on a rail that has to stop on a mark. Everything true about the cart at one instant (its position, its velocity, the friction under its wheels, the wind pushing on it) is its state, . The state is complete by definition: if you knew it exactly, and knew the physics, you could predict the next instant.
The robot never gets the state. It gets an observation, : whatever its sensors report. A wheel encoder reports position, slightly noisy and a few milliseconds late. Nothing reports the friction or the wind. The observation is a partial, delayed and noisy view of the state.
Each tick the robot emits an action, , here a motor command. A policy, , is the rule that picks it. It can be a formula, a lookup table or, in everything that follows, a neural network. It can be deterministic, , or a probability distribution you sample from, . Because one observation rarely pins down the state, practical policies read a short history of observations, .
Run a policy from a reset until something ends the attempt (a fall, a timeout, success) and you get a trajectory or episode, . Producing one is called a rollout.
There are two ways to use this loop. Open loop means planning the whole action sequence in advance and playing it back blind. Closed loop means recomputing each action from the latest observation, so that errors are noticed and corrected.
The last toggle makes a point that recurs throughout the guide. Feedback is only as good as the observation it feeds on, and every observation is late and noisy. Most of the engineering in this field is about which parts of the loop are closed, at what rate, and with which sensors (A4, A5).
MotivationWhy it’s needed
Every learning method in this guide trains a policy. What distinguishes them is what the policy observes, what action it emits, and what “good” means. Without these distinctions the vocabulary collapses. “The robot knows the friction” confuses state with observation. “The model outputs torques” names an action space. “It generalizes to new homes” is a claim about observations it has never seen (A6).
The gap between state and observation matters most. In simulation you can read the true state, including friction, masses and terrain. Much of sim-to-real transfer is about using that privileged information during training without depending on it at deployment (B6, B7).
Trade-offsVersus the alternatives
Closed-loop policythis concept
Recompute the action from fresh observations every tick.
- Corrects disturbances and model errors as they happen
- Works with an imperfect model of the world
- Needs sensing and compute at the control rate
- Feedback can oscillate or go unstable if delays or gains are wrong (B5)
Open-loop playback
Compute a trajectory in advance and execute it blind.
- No sensing and trivial compute; fine for repeatable factory motions
- Any disturbance accumulates uncorrected
- Useless when the world varies from run to run
TimelineHow we got here
Feedback control is older than electronics: Watt’s governor regulated steam engines in the 1780s, and Maxwell worked out when such loops oscillate in 1868. The state-space view (a hidden state, noisy observations, an estimator in between) arrived with Kalman in 1960. Reinforcement learning then contributed the words policy, episode and rollout (B1). Robot learning speaks both dialects.
Degrees of freedom and the floating base
A humanoid has thirty-odd motors, but its most important coordinates, where its body is in the room, have no motor at all. It can move them only by pushing on the ground.
MechanismFrom first principles
A degree of freedom (DoF) is one independent way something can move. A door has one: the hinge angle. A rigid body floating in space has six: three for position and three for orientation. Each motorized joint adds one actuated DoF.
A fixed-base arm bolted to a table is easy to reason about. Its base never moves and every DoF is a joint with a motor, so commanding seven joint angles commands everything.
A humanoid is different. A Unitree G1 has 23 to 29 actuated body joints, and up to 43 with dexterous hands. Boston Dynamics’ electric research Atlas has about 50 DoF including its grippers; the production Atlas shown in January 2026 has 56. On top of the joints, the pelvis has six DoF in the world (x, y, z, roll, pitch, yaw) with no motor attached. This is the floating base, and the full configuration is “joints + 6”.
Nothing connects the pelvis to the world except contacts. To accelerate the body the robot must push on the environment, usually through its feet, and let the reaction forces move it. Newton’s law for the center of mass makes this concrete:
The only terms the robot controls are the contact forces , and the ground limits them. It can push but not pull, and it can push sideways only as far as friction allows. Each contact force must stay inside a friction cone:
Any force outside the cone is physically unavailable, however strong the motors are.
This is why legged control is harder than arm control. A fixed arm commands every DoF directly. A humanoid must choose where and how it touches the world so that the reaction forces do what it wants, and a bad step can cost it control of its six most important coordinates.
MotivationWhy it’s needed
The floating base explains most of Parts B and C. It is why balance can’t be demonstrated by a human operator at the torque level (D1), why locomotion is learned by reinforcement learning in simulation (B2), why human motion must be checked for physical feasibility before a robot can follow it (C2), and why modern stacks put a learned whole-body controller under everything else (C6). Some layer has to own the contact forces so that the layers above can say “move the hand there” without falling over.
Trade-offsVersus the alternatives
Floating-base humanoidthis concept
Legs, torso and arms; balance through contacts.
- Goes where people go: stairs, kitchens, ladders, cars
- Can use its whole body to lean, brace and carry
- Underactuated: six unactuated DoF and constrained contact forces
- Falls are possible, dangerous and expensive
Fixed-base arm
Bolted down; every DoF is actuated.
- Simple, precise and stable; decades of mature control theory
- Reach limited to its workcell
Wheeled base with arms
A mobile manipulator on a statically stable base.
- No balance problem; cheaper and safer
- Can’t climb stairs or step over clutter; wide footprint
Quadruped
Four legs, sometimes with an arm on top.
- Statically stable with three feet down; forgiving on rough ground
- Poor at two-handed manipulation at human height
TimelineHow we got here
Walking machines were first stabilized with the zero-moment point criterion (1968), which keeps the center of pressure inside the foot. Honda’s P2 (1996) and ASIMO (2000) walked this way, carefully and flat-footed. Raibert’s hopping robots (1980s) showed a different route, dynamic balance, which Boston Dynamics carried into Atlas. The recent wave of humanoids (Unitree H1 and G1, electric Atlas, Figure 01 to 03) arrived at the same time as learned controllers that treat contacts as something to discover in simulation rather than schedule by hand.
PD control: position targets vs torque
Learned policies rarely command motor torque directly. They set a moving target for each joint, and a fast spring-and-damper loop on the motor driver does the rest.
MechanismFrom first principles
Every motor ultimately produces torque. The simplest controller that makes a joint go somewhere acts like a virtual spring and damper:
Here is the measured joint angle, is the target, and is usually zero. is a stiffness, torque per radian of error, like a spring anchored at the target. is a damping, torque per unit of velocity, like a shock absorber. This is PD control. The integral term of full PID control is usually left out for joints because it “winds up” when a limb is blocked by contact.
Seen this way, a policy that outputs isn’t commanding a position. It chooses where to anchor the spring, and the gains decide how hard the joint pulls. This is impedance control: controlling the relationship between motion and force rather than either one alone.
A worked example. Let N·m/rad. The joint is at rest, 0.1 rad short of its target, so it produces 4 N·m. If the limb carries a load that needs 6 N·m to hold, the joint settles where , which is 0.15 rad below the target. A policy trained with this controller learns to aim above where it wants the joint to end up.
Four reasons nearly every learned locomotion and whole-body policy outputs position targets:
- Disturbance rejection between ticks. The PD loop runs at about 1 kHz on the motor driver; the policy runs every 20 ms. A kick at millisecond 3 is resisted immediately, not 17 ms later.
- Latency tolerance. If the next target arrives late, the joint keeps servoing to the last one instead of going limp.
- Safe exploration. RL explores by adding noise to actions. Noise on targets makes a joint jiggle around a pose; noise on torques integrates into runaway motion.
- Hiding actuator ugliness. Friction, backlash and current limits sit inside a feedback loop, which shrinks the gap between simulation and reality (B6).
The gains are part of the action space. The same target produces different motion under different gains, so changing and after training changes the policy’s behavior.
Torque control, where the policy outputs directly, is more expressive. It can be exactly compliant or apply a precise force, and it avoids the artificial stiffness a PD loop imposes. But it needs a policy running at a very high rate, its exploration is dangerous, and the simulator must model torque tracking accurately. Figure’s S0 whole-body controller emits joint-level actuator commands at 1 kHz, which is possible only because the network is tiny, about 10 million parameters (C6).
MotivationWhy it’s needed
Position targets with a PD loop turn an unforgiving control problem into a learnable one. The PD loop supplies stiffness and damping at a rate no neural network on the robot could match, so the policy only has to decide where joints should go every 10–20 ms. That separation of timescales (A4) is what made RL-trained locomotion work on real hardware at all, and it is why the gains are among the first hyperparameters a sim-to-real engineer checks.
Trade-offsVersus the alternatives
Position targets + PDthis concept
The policy picks targets; a 1 kHz loop turns them into torque.
- Robust to latency; safe exploration; hides actuator detail
- Works with policies at 30–100 Hz
- Imposes a fixed stiffness; force control is indirect
- Gains must match between simulation and hardware
Direct torque control
The policy outputs motor torque.
- Full control of forces and compliance
- Needs kHz-rate policies and an accurate actuator model
- Random exploration throws the robot on the floor
Variable impedance
The policy outputs targets and gains.
- Can stiffen for precision and soften for contact
- Larger action space; gains interact with everything learned
Velocity control
The policy outputs joint velocities; a loop tracks them.
- Natural for arms following smooth paths
- Drifts in position; awkward for holding a pose under load
TimelineHow we got here
Proportional-integral-derivative control was analyzed in 1922 for steering ships, and Hogan’s impedance control (1985) reframed manipulation as shaping how a limb responds to force. Series elastic actuators (1995) and MIT’s proprioceptive, low-gear-ratio actuators (2017) made joint torque measurable and controllable. When deep RL reached real legged robots, Tan et al. (2018) and Hwangbo et al. (2019) both used PD position targets, and that interface became the default. Figure’s S0 (2026) is a notable move back toward direct, high-rate actuator commands.
Control frequencies
A humanoid runs several control loops at once, from a thousand times a second at the motors to once every few seconds for deciding what to do next, because physics and computation each set their own clock.
MechanismFrom first principles
Physical events happen on very different timescales. A foot strikes the ground in a few milliseconds. A step or a reach takes a few hundred. Wiping a counter takes seconds, and loading a dishwasher takes minutes.
A control loop has to run several times faster than the dynamics it shapes; a common rule of thumb is five to ten times. At the same time, computation per tick grows with model size. A forward pass through a 3-billion-parameter model on an onboard GPU takes tens to hundreds of milliseconds, while a 10-million-parameter network takes well under one. So nothing can be both big and fast, and stacks are layered by rate:
| Layer | Typical rate | Examples | Handles |
|---|---|---|---|
| Motor current and PD servo | 1–40 kHz | motor driver firmware | stiffness, damping, impacts |
| Whole-body controller | 50 Hz – 1 kHz | RL motion trackers; Figure S0 at 1 kHz | balance, contact forces |
| Visuomotor policy | 50–200 Hz | Helix S1 at 200 Hz; GR00T N1 action head at 120 Hz; π0 at 50 Hz | reaching, grasp corrections |
| Reasoning model | 1–10 Hz | Helix S2 at 7–9 Hz; GR00T N1 VLM at 10 Hz | what to do and where |
| Orchestrator | below 1 Hz | Gemini Robotics-ER agent | multi-step plans, replanning |
The pattern is consistent. The faster the loop, the smaller the model and the more local its job; the slower the loop, the bigger the model and the more it knows.
MotivationWhy it’s needed
Rates explain architecture. Dual-system designs exist because a VLM can’t close a 200 Hz loop (D13). Action chunking exists because a 50 Hz policy that thinks once per 50 steps can afford a bigger model (D3). Real-time chunking exists because inference latency is a sizable fraction of a chunk (D12). The whole-body controller exists because balance needs a millisecond-scale loop that no manipulation policy can provide (C6).
Trade-offsVersus the alternatives
Layered by ratethis concept
Several loops, each sized to its timescale.
- Every timescale gets a controller fast enough for it
- Layers can be trained on different data
- Interfaces between layers become bottlenecks and failure points
One network at the fastest rate
End to end, pixels to torques at 1 kHz.
- No hand-designed interfaces
- Needs a large model at kHz rates; infeasible on robot hardware today
One network at a slow rate
A large model directly commands joints at a few Hz.
- Simple; all knowledge in one place
- Can’t balance or react to slips; motion is jerky
Event-driven control
Replan only when something changes.
- Saves compute when nothing happens
- Hard to guarantee reaction time; still needs a fast safety loop
TimelineHow we got here
Multi-rate control is standard in aerospace and industrial robotics, where inner current loops run far faster than outer position loops. What changed recently is what sits in the slow layers: first language models used as planners (SayCan, 2022), then VLAs emitting actions at 50 Hz (π0, 2024), then explicit fast/slow splits (Helix, GR00T N1, 2025), and finally learned controllers at the fastest layer too (Figure’s S0, 2026).
Proprioception vs exteroception
Proprioception is the robot’s sense of its own body; exteroception is its sense of the world around it. A surprising amount can be done with the first alone.
MechanismFrom first principles
Proprioception is internal sensing. Joint encoders report each joint’s angle, and velocity by differentiation, usually at 1 kHz or faster. An inertial measurement unit (IMU) in the pelvis or torso reports angular velocity and linear acceleration, from which orientation relative to gravity is estimated. Motor current gives an estimate of joint torque, good on low-gear-ratio actuators and poor on high-ratio ones (I1). Foot force sensors, where present, report contact.
Exteroception is sensing the outside world: RGB cameras, depth cameras, lidar. Touch sits between the two. Fingertip tactile sensors measure contact the way cameras can’t: Figure 03’s fingertips register forces as small as 3 grams (G8).
The two families differ in more than subject. Proprioception is low-dimensional (a few dozen numbers), fast, and nearly noise-free. Cameras produce hundreds of thousands of numbers per frame at 30–60 Hz, and those numbers are hard to interpret.
“Blind” locomotion uses proprioception alone, and it works surprisingly well. A short history of joint states and IMU readings implicitly reveals the terrain and the robot’s own dynamics. If a foot touches down earlier than commanded, the ground is higher than expected; if a joint lags its target, the load is heavier. A policy with memory can infer these hidden quantities without seeing anything (B6).
MotivationWhy it’s needed
Manipulation needs exteroception: you can’t find a mug by feel from across the room. But every exteroceptive channel adds cost. Pixels are high-dimensional, simulated images differ from real ones, and camera latency is long. That is why locomotion was mastered blind first (B5), why perception is often added as a compact height map rather than raw pixels, and why tactile sensing is valuable yet hard to share across robots (G8).
Trade-offsVersus the alternatives
Proprioception onlythis concept
Encoders, IMU and motor currents.
- Fast, low-dimensional and easy to simulate accurately
- Robust across lighting and weather
- Can’t see a step before touching it, or find an object
Proprioception + height map
Terrain heights from depth or lidar, fused with a belief state.
- Anticipates stairs and gaps
- Maps are noisy; the policy must learn when to distrust them
End-to-end vision
Raw camera images into the policy.
- Needed for manipulation and semantics
- High-dimensional, slow, and hard to simulate faithfully
Touch
Tactile skins and fingertips.
- Senses slip and grip force directly
- No standard sensor or data format; little shared data
TimelineHow we got here
State estimation from inertial sensors goes back to the Kalman filter (1960). In learned locomotion, ETH Zurich’s blind ANYmal policy (2020) and Oregon State’s blind stair-climbing Cassie (2021) showed how far proprioception goes; Miki et al. (2022) then fused noisy height maps into the same kind of controller. In manipulation the trend runs the other way, toward more exteroception: palm cameras and fingertip tactile sensors on Figure 03 (2025), fed directly into the visuomotor policy by Helix 02 (2026).
Zero-shot, seen, unseen, generalization
A policy generalizes when it succeeds on situations it was not trained on. Whether a situation counts as new is harder to pin down than it sounds.
MechanismFrom first principles
A learned policy is fitted to a training distribution: the range of scenes, objects, instructions and physical conditions in its data. On situations that resemble that data (in-distribution) it interpolates between things it has seen. On situations outside it (out-of-distribution) it has to extrapolate, and learned models extrapolate badly.
Papers describe test conditions as seen (in the training data) or unseen (absent from it), usually along a named axis:
- Visual: new lighting, backgrounds, camera positions.
- Object: new instances or categories of objects.
- Scene: new rooms, homes or factories.
- Semantic: new phrasings or instructions (“the extinct animal”).
- Task: new combinations of known skills, or new skills.
- Embodiment: a robot body the model never trained on (D9).
Zero-shot means succeeding with no task-specific training at all. Few-shot means adapting from a handful of examples or a short fine-tune (D11).
The figure shows the core difficulty. A model can have near-zero error on everything like its training data and be arbitrarily wrong a small step beyond it. Neural networks behave better than high-degree polynomials, but the principle holds: their guarantees stop at the edge of the data.
Defining “unseen” is getting harder as datasets grow. When a model has been pretrained on the internet and hundreds of thousands of hours of robot data, a “novel” kitchen may still resemble something in training. The π0.7 authors say as much. Careful papers therefore report which axes changed, and hold out whole environments rather than individual episodes (H1).
MotivationWhy it’s needed
Generalization is the reason for foundation models. A policy that works only in the lab where its demos were collected is a demo, not a product. Every data strategy in Part G is ultimately about widening the training distribution cheaply, and every claim you read (“56% in unseen homes”, “90% on seen tasks”) is only meaningful once you know what was held out.
Trade-offsVersus the alternatives
Approaches differ in how they buy generalization:
More diverse data
Collect across many scenes and objects (G7).
- The most reliable lever measured so far
- Expensive per scene for robot data
Pretrained representations
Start from web-scale vision-language or video models (D6, F5).
- Semantic generalization to new objects and instructions for free
- Doesn’t supply physical skill
Randomization
Vary simulated physics and visuals (B5).
- Cheap coverage of physical variation
- Only covers variation the simulator can express
TimelineHow we got here
Zero-shot transfer became a headline capability with CLIP (2021), which classified images into categories it was never explicitly trained on. RT-2 (2023) carried web semantics into robot control. Since then the frontier has been physical and environmental generalization: Lin et al. (2024) measured that diversity of environments, not number of demos, drives it; π0.5 (2025) and Helix 2.5 (2026) report success in homes their robots had never entered.
Neural-network vocabulary
Parameters, embeddings, tokens, attention, transformers, ViTs, DiTs, losses, fine-tuning, LoRA and distillation: the words every later section uses.
MechanismFrom first principles
Parameters are a network’s learned numbers (weights and biases). “A 3B model” has three billion of them.
An embedding or latent is a vector of numbers that represents something, such as an image patch, a word or a goal, in a space where similar things end up near each other. Networks do their work on embeddings.
A token is one item in a sequence that a transformer processes: a word piece, an image patch, a discretized action (D8). Everything the model reads or writes is converted to tokens and embedded.
Attention lets each token build its representation by mixing information from other tokens. Each token computes a query , a key and a value (three linear projections of its embedding). Token weights token by how well its query matches ’s key:
A transformer stacks layers of attention and small per-token MLPs. Because every token can attend to every other, cost grows with the square of sequence length, which is why token counts matter for robot latency (D6).
A ViT (vision transformer) chops an image into patches, typically 14 or 16 pixels square, embeds each patch as a token, and runs a transformer over them. A 224×224 image in 14-pixel patches becomes 16 × 16 = 256 tokens. A DiT (diffusion transformer) is a transformer used as the denoiser inside a diffusion or flow model (D4).
The loss is the number training minimizes. Mean squared error (MSE) is used for regression, predicting continuous values. Cross-entropy is used for classification and next-token prediction. Validation loss is the loss on held-out data, used to measure generalization without running the robot; it is only loosely coupled to real task success (G6).
Fine-tuning continues training a pretrained model on new data. Freezing locks some weights so they don’t change. LoRA (low-rank adaptation) freezes a weight matrix and learns a small correction , where and are thin matrices of rank ; for a matrix it trains numbers instead of (D11). Distillation trains a student model to reproduce a teacher’s outputs (B7, E6).
MotivationWhy it’s needed
These are the parts every architecture in this guide is assembled from. A VLA is a ViT plus a language model plus an action head (D7); a teacher-student pipeline is distillation (B7); a post-trained robot model is often a LoRA fine-tune (D11). Knowing that attention is quadratic in tokens explains why action tokenizers (D8) and image resolution choices matter for latency.
Trade-offsVersus the alternatives
Robotics still uses every major architecture, each where it fits:
MLP
Stacked fully connected layers.
- Tiny and fast; the default for low-level controllers
- No memory, no spatial structure
Recurrent nets (GRU, LSTM)
Carry a hidden state through time.
- Cheap memory for history-dependent control (B6)
- Hard to train on long horizons; sequential
Convolutional nets
Local filters over images.
- Efficient visual features; strong inductive bias
- Less flexible than attention at scale
Transformersthis concept
Attention over tokens of any modality.
- Scale well; mix images, text, state and actions in one sequence
- Quadratic cost in sequence length; data-hungry
TimelineHow we got here
The transformer (2017) was built for translation; the ViT (2020) showed images could be read the same way, and diffusion transformers (2022) put the architecture inside generative models. LoRA (2021) made adapting giant models cheap. Distillation is older than all of these: model compression (2006) and Hinton’s knowledge distillation (2015) trained small students to mimic big teachers, the same move robot learning uses for teacher-student policies.