Files: MD · PDF

Contents

Dexterous Humanoid Hands: Landscape, Benchmarks & How to Reproduce SOTA

Oct 2, 2026 · @julian gilyadov

The status quo to beat is an open, low-cost hand (ORCA or LEAP) trained by glove teleop, GPU sim-to-real RL, or pretraining on human video. The best real results now come from human-data pretraining, while real in-hand reorientation on cheap hands still manages only a few consecutive goals.

To know whether our work improves on this, reproduce four anchors on one hand in about eight weeks: simulated reorientation, real reorientation, glove-teleop imitation learning, and POMDAR plus a human-data baseline. Freeze them in the baseline sheet, then test our data, simulator or environment on the metric that matches it.

The biggest openings: nobody has published autonomous POMDAR scores or ORCA reorientation numbers, and real reorientation on low-cost hands lags simulation badly.

Hardware landscape

Labs train on cheap open hands (LEAP, RUKA, ORCA: about $200–2,000 in parts), while 20–22 DoF tactile hands (Sharpa, Wuji, Tesla) set the 2026 frontier but publish few comparable benchmarks.

Research hands you can buy or build (sorted by price)

HandReleasedActuated DoFActuationTactilePrice (USD)Known for
LEAP Hand v2 (CMU, 2025)Jun 2025 (RSS paper)8Hybrid rigid-soft, servosNone reported~200 partsUnder 2 h assembly
RUKA (NYU, 2025)Apr 2025 (arXiv)11 (15 joints)Tendon, forearm motorsNone~1,300 partsLearned controllers from MANUS-glove data; Kapandji 10/10; 29 of 33 GRASP-taxonomy grasps
LEAP Hand v1 (CMU, 2023)Jul 2023 (RSS paper)16 (4 fingers)Servos in jointsNone~2,000 parts; 2,066 kitMost-used low-cost research hand
ORCA (ETH Zurich, 2025)v1: Apr 2025 (arXiv); v2 line (lite, standard, Touch): Mar 15, 2026; v2.1: no public date found17 incl. wristTendon, forearm motorsFingertip FSR; Touch: 351 taxels~2,000 parts; 3,500+ assembledPublished reliability data, LeRobot stack
Wuji HandSep 18, 2025 (Hand 2 teased Jun 2026)20Micro-actuator per jointNone standard16,000 US list580 g, 15 N fingertip, 300k+ grasp-cycle rating
Inspire RH56DFX / BFX2022 (DFX); BFX date not published6 active (12 joints)Linear servosOptional fingertip force20,599–21,599 US listCommon on Unitree humanoids
Allegro Hand2012; V4 2018, V5 202416 (4 fingers)Motors in joints, CAN busNone standard21,450 US listMost-cited hand; HORA and DexPoint baselines
Sharpa WaveOct 2025 (first shipments)22Proprietary1,000+ taxels per fingertip at 180 HzNot disclosed30 N fingertip force; mass production claimed
Shadow Dexterous Hand200520 (24 joints)TendonExtensive> 100,000 CHFOpenAI Dactyl and Rubik's cube

US list prices for LEAP v1, Wuji, Inspire and Allegro are one reseller's (Robotics Center, July 2026); Shadow's cost is from the ORCA paper.

Release dates: research hands use the first paper date; commercial hands use first sale or first shipment (Wuji, Inspire, Allegro, Sharpa, Shadow, ORCA v2 line). ORCA v2.1 appears in no public release note or the orca_core repository, so its date is open.

Humanoid OEM hands (not sold separately; reported specs as of April 2026, Wikipedia)

Robot handReleasedDoFActuationSensing
Tesla Optimus Gen 322-DoF hand shown Oct 2024; full Gen 3 robot not revealed as of Apr 202622Tendon; 25 actuators in the forearmVision-first; tactile little disclosed
Sanctuary Phoenix21-DoF hand Dec 2024; Phoenix Gen 7 Apr 202420–21HydraulicMicro-barometer pads, 5 mN sensitivity
Figure 02 (Helix 02)Aug 6, 202416Electric unit per fingerFingertip tactile (~3 g), palm cameras
Fourier GR-2Sep 30, 202412Electric6 tactile arrays per hand
Unitree H2 (Dex5 option)Oct 20, 202510–12ElectricOptional tactile
Apptronik Apollo (PSYONIC Ability)Apollo Aug 2023; Ability Hand 2021Low; prosthesis-derivedElectricMulti-touch

DoF count is a weak proxy for dexterity. A May 2026 kinematic metric across 16 hands (KaRMA) ranked LEAP highest and Ability lowest, and found DoF alone does not predict reachable workspace.

Release dates are first public unveilings (Optimus, Sanctuary hand, Phoenix Gen 7, Figure 02, GR-2, H2, Apollo, Ability Hand). Benchmark and dataset dates elsewhere are first paper or release dates (Elliott & Connolly, Isaac Lab, EgoDex, ActionNet, RoboMind, RealDex, BrainCo Revo2).

ORCA deep dive

ORCA is the best-documented low-cost anthropomorphic research hand: 17 DoF, tendon-driven, under 2,000 CHF in parts, with published reliability, accuracy and learning results (ORCA paper, IROS 2025). Since June 2026 it also has an open learning stack that plugs into Hugging Face's LeRobot (ORCA platform paper).

Design (v1)

  • Kinematics: 16 finger joints plus 1 wrist joint. Fingers 2–5 have MCP, PIP and abduction joints (no DIP); the thumb adds a CMC joint and sits at 15° supination for opposition.
  • Actuation: each joint is pulled by a flexor/extensor pair of 0.4 mm braided nylon tendons over metal pins. Motors sit in a forearm "tower" with fans; the wrist is belt-driven with 60° of flexion and extension.
  • Reliability features: pin joints that pop out instead of breaking, ratchet spools for re-tensioning in seconds, and auto-calibration that drives each joint to its stops to fit a motor-to-joint ratio without joint encoders.
  • Sensing: binary force-sensing resistors under silicone skin on all five fingertips.

Published data points

MetricResultConditions
Build< 2,000 CHF materials, < 8 h assembly by one personDIY, 3D-printed PLA
Continuous grasping2,250 grasp cycles in 2.5 h, no failureGrasp every 4 s, wrist cycle every 16 s; stopped by choice
Durability claim> 10,000 cycles (~20 h) without hardware failureAbstract figure
Joint trackingAccuracy similar to LEAP Hand; mean latency < 0.2 s2 Hz and 5 Hz sine, AprilTags filmed at 60 fps; latency mostly software
Sim-to-real RLZero-shot tennis-ball reorientation about a given axisIsaacGymEnvs, 4,096 parallel envs, A2C, ~1 h training with domain randomization
Imitation learning214 demos (~2.5 h) gave a policy that ran 7 h 17 min (~2,000 grasps) with no hardware interventionRokoko-glove teleop, Franka arm, 3 cameras, diffusion transformer, 500 epochs ≈ 4 h on one RTX 4090
IL ablationMasked-cube input beat raw RGB60 trials, 10 per table sub-area
Teleop skillsCube stacking, jar-cap twist, fidget-spinner spin, writing "Hello", pouring 50 mlQualitative, gloves
Tactile limitsDetects 0.05 N; skin wear after ~2,000–4,000 cycles; sensor wires snapped after ~4,500–7,000v1 binary FSR design

Buying options (Jul–Oct 2026)

OptionPrice (USD)Contents
BOM kit643Tendons, connectors; no motors, prints or silicone
Assembly kit with motors4,234 per handDynamixel motors, tested prints, cast skin, fans
Fully assembled v1 (2025)5,929 per handTested, hard case, spares
Assembled ORCA Hand, current listingfrom 3,500 (Feetech) or 4,500 (Dynamixel)Manufacturer-confirmed Jul 31, 2026
ORCA Hand Touch (Mar 2026)6,100–7,100351 Hall-effect taxels, 6D force per taxel, ~1.25 kg
9-DoF lite (Mar 2026)BOM < 900Coupled tendons, self-build only

The seller is ORCA Dexterity, Inc., a Delaware company with its team in Zurich; no grip-force or payload figure is published (RoboZaps record).

Software

  • orca_core: MIT-licensed Python controller with tension, calibrate and neutral scripts. It auto-detects Dynamixel or Feetech motors and ships v1 and v2 hand configs (~590 stars, 150 forks).
  • Teleop: Rokoko gloves and Apple Vision Pro are supported out of the box; the retargeting code is open (shop page).
  • ORCA learning stack (June 2026): one interface for control, simulation, consumer-device teleop and retargeting. Its reference workflow is VR-headset teleop of in-hand reorientation, a LeRobot-trained policy, and a reproducible evaluation setup.

What ORCA has not published: grip force, payload, YCB-style grasp success rates, or consecutive-success counts for reorientation. Those gaps are exactly where a new data or simulation product can show a measurable gain.

How dexterous hands get their skills

Five training routes dominate in 2026, and the fastest-moving one is human data: egocentric video or wearable capture, converted into robot trajectories and co-trained with a few robot demos.

RouteHow it worksTypical costMain limitation
Teleop + imitation learningAn operator drives the hand with gloves or a VR headset; a policy (Diffusion Policy, ACT, π0-style VLA) clones the demos, e.g. ORCA (Apr 2025)50–200+ robot demos per taskSlow and tied to one hand; tracking quality caps data quality (gloves beat headsets under occlusion)
Sim-to-real RLTrain in a GPU simulator with domain randomization, deploy zero-shot, e.g. MuJoCo Playground (Feb 2025)Minutes to days of GPU timeNeeds object-pose tracking and reward design; real results trail simulation
Human video to robotRetarget hand poses from egocentric video into robot trajectories, then pretrain, e.g. UniDex (Mar 2026), EgoScale (Feb 2026)Thousands of hours of videoKinematic and visual gaps; still needs some robot data
Human–robot co-trainingWearable rigs (gloves, head cameras) record humans; a few robot demos anchor the embodiment, e.g. DexWild (May 2025)Human demos collect several times faster than robot demosSuccess collapses without any robot demos
World models and dexterous VLAsPretrain dynamics or policies on mixed human and robot data, then plan or fine-tune per hand, e.g. DexWM (Dec 2025)Large pretraining, small fine-tuneCross-hand transfer is still weak

Two caveats recur across routes: human data alone does not finish the job, and policies rarely transfer across hands without retraining. The State of the art section below gives the numbers.

Where hands are used today

  • Research benchmarks: in-hand reorientation, grasping, tool use (scissors, spray bottles, kettles) and in-hand assembly.
  • Humanoid demos: Sharpa's robot assembled a paper windmill in 30+ autonomous steps at CES 2026 (report). OEMs publish videos far more often than success rates.
  • Data collection: teleop rigs (ORCA supports Rokoko gloves and Apple Vision Pro) feed imitation-learning datasets.
  • Prosthetics-derived hands: PSYONIC Ability (2021) and BrainCo Revo2 (Sep 2025) hands carry prosthetic designs onto humanoids.

Benchmarks and metrics

There is no single dexterity leaderboard. A credible claim stacks three levels: a hand-level benchmark (POMDAR), a simulation suite (MuJoCo Playground), and a real-world policy protocol with fixed trial counts and reported variance.

Hand-level benchmarks (what the hand plus controller can do)

BenchmarkMeasuresProtocol and scoreStatus
POMDAR (ETH Zurich, Apr 2026)12 in-hand manipulation and 6 grasp tasks from the Elliott & Connolly and GRASP taxonomies3D-printed scaffolds; score = 0.8 × correctness + 0.2 × speed vs a human baseline; 20 trials per task; MuJoCo twinORCA scored by teleop only (1,140 trajectories, ~25 h); gloves beat Apple Vision Pro on occluded tasks. No autonomous-policy scores yet
Elliott & Connolly benchmark (CMU, Jul 2021)13 in-hand patterns using the digits onlyYCB objects, visual trackingNo grasp tasks
HD-marks (2020)50 tasks: GRASP grasps, Kapandji thumb postures, in-hand axesMostly binary successBroad, weak cross-lab comparability
Kapandji test (1986) + GRASP taxonomy (2016)Thumb opposition (0–10) and 33 grasp typesStatic posturesRUKA: 10/10 and 29 of 33
KaRMA (May 2026)Kinematic workspace for fine manipulationComputed from hand models16 hands ranked
DexBench (RLWRLD, Jun 2026)18 atomic industrial tasks in ~80 cases, placed on object-complexity and dexterity-regime axesReal-world, customer-derivedEndorsed by Lotte, SK Telecom, CJ Logistics and others

The Elliott & Connolly and HD-marks rows are as summarized in the POMDAR paper.

Simulation suites

SuiteContentsHeadline metric
MuJoCo Playground (Feb 2025)LeapCubeReorient and LeapCubeRotateZAxis; open source, pip-installableConsecutive successes; ~35 min to train on one RTX 4090
Isaac Gym (2021) / Isaac Lab (Jun 2024)GPU RL used by DeXtreme and by ORCA's ball-rotation policyConsecutive successes
Bench2Dex (Sep 2026)26 bimanual visuo-tactile tasks, 12 hands, ~1.3K teleop demosSuccess, with mandatory seeds, episode counts and termination rules
DexVerse (Jul 2026)Modular multi-task, multi-embodiment suiteNot reviewed here
LabDex (Aug 2026)Hierarchical chemistry-lab skills, simulation and real robotSkill-to-long-horizon scaling
In-hand assembly (Sep 2026)Two-part assembly inside one hand, Isaac SimGoal-reaching error by hand morphology

Real-world policy metrics worth reporting

MetricUsed byDefinition
Consecutive successesDeXtreme, MuJoCo PlaygroundGoals reached before a drop; median and mean over ≥ 10 trials
Task progress + final successUniDexMean stage completion and full-task success over 20 trials
Unseen-scene successDexWildSuccess in environments absent from training
Zero-shot cross-hand transferUniDexTask progress on a hand never trained on
Human–robot exchange rateUniDexHuman demos needed to replace one robot demo (≈ 2:1)
Autonomy hoursORCAHours or cycles without hardware intervention
Contact fidelityTactiDex (Jul 2026)Agreement of contacts and forces with human demos

Datasets

DatasetSizeContent
EgoDex dataset (Apple, May 2025)829 h of 1080p videoEgocentric human manipulation with hand-pose annotations
UniDex-Dataset (Mar 2026)52K trajectories, 9M frames, 8 robot handsRetargeted from H2O, HOI4D, HOT3D and TACO
ActionNet (Fourier, Mar 2025)30K trajectories, 2 handsDexterous bimanual teleop
RoboMind (Dec 2024)19K trajectories, 1 handMulti-embodiment teleop
RealDex (Feb 2024)2K trajectories, 2 handsHuman-like grasping

ActionNet, RoboMind and RealDex sizes are from the UniDex paper's comparison table.

State of the art

The strongest 2026 results come from human-data pretraining: UniDex-VLA doubled π0's real tool-use success, and EgoScale showed policy quality scales predictably with hours of egocentric video. Real in-hand reorientation on low-cost hands remains weak.

DateWorkHandResult
Sep 2026In-hand assemblySharpa Wave; Wuji, Allegro, XHand in simSharpa and Wuji performed best; Allegro and XHand failed deep insertion
Jul 2026UHASLEAPReal cube reorientation averaged 2.0 consecutive goals; multi-hand training scored lower than single-hand
Jun 2026ORCA learning stackORCAFirst open, LeRobot-native loop: VR teleop, policy training, reproducible evaluation (no headline number)
Mar 2026UniDex-VLAInspire, Wuji81% task progress and 76% final success vs π0 at 38% and 35%, with 50 demos per task; zero-shot transfer reached 40% on Wuji and 60% on Oymotion
Feb 2026EgoScale22-DoF handPretrained on 20,854 h of human video with a log-linear scaling law; average success rate up 54% over no pretraining after ~50 h aligned human and 4 h robot data
Jan 2026Sharpa NorthSharpa WavePaper windmill assembled in 30+ autonomous steps (demo, no success rate)
Dec 2025DexWMAllegro10 of 12 real grasps (~83%) without robot fine-tuning
May 2025DexWildMultiple68.5% success in unseen environments, ~4× robot-only; 5.8× better cross-embodiment generalization
Feb 2025MuJoCo PlaygroundLEAPZero-shot cube reorientation: median 3.5, mean 7.1 consecutive goals over 10 trials (best 27)
Oct 2022DeXtremeAllegroVision-based reorientation: best run 112 consecutive goals; mean 23.1 when capped at 50; 60 h of training

Real in-hand reorientation varies about 10× by setup: DeXtreme averaged 23 goals on Allegro, while recent low-cost LEAP results average 2–7. Humanoid OEMs (Tesla, Figure, Sanctuary, Sharpa) publish demos rather than success rates, so their capability cannot be benchmarked from outside.

Reproduction plan

Reproduce four anchors in about eight weeks with one ORCA or LEAP hand, one arm and one RTX 4090-class GPU: simulated reorientation, real reorientation, glove-teleop imitation learning, and a human-data baseline. Freeze the protocol before testing our own work.

Diagram: reproduction roadmap · 4 phases, 4 gates (shown in the PDF version).

Each gate must pass before the next phase starts; the frozen sheet is what our own work is measured against.

Phase 0: simulation only (week 1, no hardware)

  1. pip install playground, then train LeapCubeReorient with Brax PPO using the published settings (100M steps, 8,192 envs). Expect ~35 min on one RTX 4090; run 3 seeds (MuJoCo Playground).
  2. Log simulated consecutive goals at 0.1 rad tolerance and wall-clock time to a fixed reward.
  3. Load ORCA in MuJoCo through the ORCA learning stack and run the same task, so later comparisons are hand-matched.
  4. Install POMDAR's MuJoCo version (project page) and teleoperate a few tasks to validate the pipeline.

Gate 0: reward curves and wall-clock match the paper within about 20%.

Phase 1: real in-hand reorientation (weeks 2–3, no arm needed)

  1. Build the hand. ORCA assembles in under 8 h; then run scripts/tension.py, scripts/calibrate.py and scripts/neutral.py from orca_core. LEAP v1 takes about 3 h.
  2. Copy the Playground rig: palm tilted 20° down on an 80/20 frame, one RealSense D415 overhead, a 7 cm cube, policy at 20 Hz, and DeXtreme's cube-pose detector or AprilTags.
  3. Deploy zero-shot. Run 10 trials, ending each when the cube drops or stalls for 30 s. Report every trial plus median and mean consecutive goals at 0.4 rad tolerance.

Gate 1: LEAP matches or beats the published LEAP numbers in State of the art.

Phase 2: teleop and imitation learning (weeks 3–6, arm required)

  1. Mount the hand on a 7-DoF arm (Franka-class; ORCA uses an ISO 9409-1 flange) with two external cameras and one wrist camera.
  2. Teleoperate with Rokoko gloves, calibrating retargeting each session. Apple Vision Pro also works but scores lower when fingers are occluded.
  3. Collect ~200 demos of ORCA's self-resetting cube pick-and-place (about 2.5 h), plus 50 demos each for two UniDex-style tool tasks, such as kettle pouring and spray-bottle pressing.
  4. Train with LeRobot: Diffusion Policy and ACT as baselines, plus a π0-class VLA fine-tune. Budget ~4 h per policy on one RTX 4090.
  5. Evaluate pick-and-place over 60 trials across 6 table zones (ORCA protocol), and each tool task over 20 trials with stage-wise progress (UniDex protocol). Add one multi-hour autonomy run.

Gate 2: multi-hour autonomy without hardware intervention, and tool-task results within about 10 points of the published baselines.

Phase 3: hand-level score and human-data baseline (weeks 6–8)

  1. Print the POMDAR rig. Score teleop first (20 trials per task) to calibrate against ETH's published ORCA plots, then score autonomous policies, which no one has published yet.
  2. Build a UniDex-Cap-style capture rig (Apple Vision Pro plus a RealSense L515 on a printed mount), or start from public EgoDex or UniDex-Dataset data.
  3. Co-train human and robot demos on one task. Sweep the human-demo count at fixed robot-demo counts to measure the exchange rate.

Gate 3: a frozen baseline sheet in which every published number is reproduced or its gap explained.

What you need

  • One ORCA (17 DoF, tendon) or LEAP v1 (16 DoF, servos-in-joints) hand; buying both gives a morphology check. Prices are in the sections above.
  • A 7-DoF arm with an ISO 9409-1 flange for Phases 2–3; Phase 1 needs only a frame.
  • One RTX 4090-class GPU; every baseline above trains on one.
  • RealSense D415 or D435 for in-hand tracking; L515 for UniDex-style RGB-D capture.
  • Rokoko or MANUS gloves; Apple Vision Pro optional.
  • Open question: arm, glove and camera prices are not covered here.

Evaluation protocol: proving our work beats the status quo

A claim holds only if, on the frozen protocol above, our data, simulator or environment moves a reproduced baseline number at equal or lower cost, with enough trials to separate signal from noise.

Pick the metric that matches the product

If our work isBaseline to beatPrimary metricWin condition
Glove or wearable dataDexWild co-training; UniDex-CapHuman demos needed per robot demo; unseen-scene successFewer human demos per robot demo than UniDex's ratio, or higher unseen-scene success at equal robot demos
Egocentric video dataEgoScale scaling curve; EgoDex; UniDex-DatasetReal task success after pretraining on matched hours; loss-vs-hours slopeEqual success from fewer hours, or a steeper scaling curve
Simulator or sim-to-real pipelineMuJoCo Playground LEAP; DeXtremeReal consecutive goals; sim-to-real drop; GPU-hoursHigher real median and mean at equal or fewer GPU-hours
Training or evaluation environmentsPOMDAR-sim; Bench2DexRank agreement between sim and real scores across at least 5 policiesHigher rank correlation with real outcomes than existing suites

Rules for a fair comparison

  • Pre-register tasks, objects, success definitions and trial counts before any run.
  • Change one variable: same hand, arm, cameras, policy architecture and compute; only the data or environment differs.
  • Randomize and blind: interleave policies A and B in random order, and keep the person resetting scenes unaware of which is running.
  • Size trials for the effect: 20 trials leave about ±22 points of 95% uncertainty at 50% success. Separating 50% from 70% with 80% power takes about 90 trials per arm.
  • Report distributions: per-trial values plus median and mean, because consecutive-goal metrics are heavy-tailed.
  • Split in-distribution results from held-out objects, scenes and hands.
  • Normalize by cost: report success per hour of data collection and per GPU-hour, since human demos collect faster than robot demos.
  • Train at least 3 seeds, and publish seeds, episode counts, termination rules and horizons, as Bench2Dex requires.

Baseline sheet

MeasurePublished referenceReproducedOursΔ (95% CI)Status
LEAP real reorientation, median / mean consecutive goalsMuJoCo Playground (2025)Not started
ORCA real reorientation, median / mean consecutive goalsNone publishedNot started
ORCA pick-and-place success, 60 trialsORCA (2025)Not started
Autonomy hours without hardware interventionORCA (2025)Not started
Tool-task progress and final success, 20 trials per taskUniDex (2026)Not started
POMDAR score, teleopPOMDAR (2026)Not started
POMDAR score, autonomous policyNone publishedNot started
Human demos per robot demoUniDex (2026)Not started
GPU-hours to train reorientationMuJoCo Playground (2025)Not started

The two "None published" rows are open territory: the first credible autonomous POMDAR and ORCA reorientation numbers would themselves define the status quo.

Sources

Opened in full:

Cited from search excerpts of the source pages: