> Working notes for *Open-Source Humanoid Robotics Landscape (Oct 2026)*, compiled Oct 2, 2026. Not fact-checked line by line: where these notes and the report disagree, trust the report. See README.md.

# Open-source robot-learning tooling, open datasets and benchmarks: the data-economy slice (as of 2026-10-02)

## How this was compiled. Read this first.

- **Sources and dates.** Everything was fetched on 2026-10-02.
  - GitHub, PyPI and crates numbers come mostly from the GitHits package/repo index (refreshed 2026-09-15 to 2026-10-02), because direct GitHub and PyPI pages were rate-limited.
  - Hugging Face (HF) numbers come from HF pages opened today. HF "downloads last month" is a rolling file-request counter, not a count of unique users, and it is inflated for many-file or streamed datasets. Use it only to compare datasets with each other.
- **Tool failures (not retried):**
  - **Rate-limited (HTTP 429) early in the session:**
    - github.com/huggingface/lerobot and /releases
    - pypi.org/project/lerobot
    - HF blog posts for LeRobot v0.5.0 and v0.6.0
    - HF docs page /docs/lerobot/lerobot-dataset-v3
    - techtimes.com article "LeRobot Hub surpasses 58,000 datasets"
    - letsdatascience LeRobot 0.6 article
    - roboticsandautomationnews NVIDIA×HF LeRobot article
    - arxiv.org/abs/2505.11709
    - egxodata.com "Robotics Data Release Tracker 2026"
    - behavior.stanford.edu 2025 leaderboard
  - **Blocked:** dexset.ai (robots.txt fetch timed out).
  - **JavaScript-rendered pages with no data returned:** robochallenge.ai, robo-arena.github.io, the RoboTwin leaderboard, the BEHAVIOR 2026 HF leaderboard Space, and xdof.ai.
  - **Budgets exhausted:** the shared WebSearch budget (200/session) and then the WebFetch budget (800/h) ran out near the end. RoboWorld (arXiv 2607.01060) and RobotArena∞ were not opened.
- **Labels used below:**
  - "(headline only)" means the fact comes from a search-result title, not an opened page.
  - "(self-reported)" means a vendor or press claim that I did not verify independently.

---

## Executive summary (for a founder supplying robot-learning data)

1. **LeRobot is now the de facto format and toolchain.**
   - HF lists **77,473 datasets tagged LeRobot** (today). A crawl in Sep 2025 found 16,065 public ones.
   - LeRobotDataset **v3.0** shipped with lerobot ≥0.4.0 (Oct 2025). New large releases use it: BEHAVIOR-2026 demos (3.27 TB), EgoSuite-Open100K and RoboCerebra.
   - Most of the volume is tiny hobby data:
     - median of 10 episodes per dataset;
     - 43% of datasets have 1–5 episodes;
     - SO-100/SO-101/Koch/LeKiwi arms make up more than half;
     - 91% are single-task.
   - Quality is weak: one audit found **81% of 100 public datasets had issues**.
2. **2026 is the year of open egocentric human data.**
   - Build AI Egocentric-100K: 100,405 h, Apache-2.0, but video only at 456×256.
   - Lightwheel EgoSuite-Open100K: 10k h live with 100k planned, hand pose, LeRobot v3 plus MCAP, and commercial training allowed.
   - EgoVerse: 1,362 h.
   - Ant Group Open-AoE: about 2,000 h from smartphones, MANO hand poses, CC BY 4.0.
   - Human Archive HA-Ego-500: 500 h, densely annotated.
   - Vendors themselves say raw footage is turning into a commodity: about $2–5/h resale, against $15–40/h list prices.
3. **The large real-robot open datasets are mostly non-commercial.**
   - Non-commercial: AgiBot World (1.0M trajectories, 2,976 h), Galaxea (500 h), Fourier ActionNet (140 h), EgoDex (829 h), LAFAN1, Motion-X, AMASS.
   - Commercially usable sets are older (OXE, BridgeData V2), simulated (NVIDIA GR00T-sim 325k trajectories, Axis 50k), or gated (RoboMIND, Apache-2.0).
4. **Evaluation is moving from saturated simulation to real-robot, distributed and long-horizon tests.**
   - LIBERO is saturated at 96–98%.
   - The BEHAVIOR 2025 winner scored a q-score of only 26%.
   - RoboChallenge's best is about 62% on Table30 V2 (per press).
   - Other real-robot efforts: RoboArena, the AgiBot World Challenge (526 teams) and ArmnetBench ($359 SO-101 cells).
   - Evaluation as a service is emerging as a niche (Cortex AI, RoboChallenge).
5. **Money is flowing into data suppliers.**
   - Raises: XDOF $70M; Mecka about $60–68M; Genesis AI $105M seed (tactile data glove); Encord $60M Series C; Config $27M; Midcentury $15M (headline only); Axis $12M; Human Archive $8.2M.
   - Vendor-published prices: **$15–60 per raw hour** (egocentric or teleoperation) and **$80–150/h for full humanoid multi-sensor** capture.
   - Collector pay in India: **$1–4.20/h**.

---

# Part A — Tooling and middleware

### LeRobot (Hugging Face): end-to-end PyTorch robot-learning library, dataset format and Hub ecosystem
- **Status: ACTIVE.**
  - Releases: v0.6.1 (2026-08-03), v0.6.0 (2026-07-06), v0.5.1 (2026-04-07), v0.5.0 (2026-03-09), v0.4.4 (2026-02-27), v0.4.3 (2026-01-22), v0.4.0 (2025-10-23).
  - Repo last pushed 2026-10-02.
- **Who:**
  - Hugging Face robotics team (Paris).
  - **Rémi Cadène** (original lead) has left. He is CEO of **UMA** (Paris), building the "Northstar" humanoid. His co-founders are Simon Alibert (CTO, LeRobot co-founder), Robert Knight (Chief Robot Officer, SO-100 arm designer) and Pierre Sermanet (CSO, ex-DeepMind). Thomas Wolf is listed as an adviser to UMA (TNW, 2026-07-07). UMA's seed round size is unconfirmed (about $40M was reportedly sought).
  - Release notes are now dominated by @imstevenpmwork (who cuts the releases), @CarolinePascal, @pkooij, @Maximellerbach, @s1lent4gnt, @nicolas-rabault, @HaomingSong and @AdilZouitine.
  - ICLR 2026 paper authors include Cadène, Alibert, Capuano, Aractingi, Zouitine, Kooijmans, Pascal, Palma, Shukor, Aubakirova, Lhoest, Gallouédec and Wolf.
- **Links:** GitHub `huggingface/lerobot`; docs huggingface.co/docs/lerobot; HF org `lerobot`; paper arXiv **2602.22818** (ICLR 2026).
- **Stats (today):**
  - **27k stars, 5.7k forks, 976 open issues.**
  - **PyPI: 238k downloads/month**, 12 versions published.
  - Contributor count not verified (GitHub page rate-limited).
- **License:** Apache-2.0.
- **What it is, in detail:**
  - **Unified Robot/Teleoperator/Camera classes.**
    - Native hardware: SO100/SO101, LeKiwi, Koch, HopeJR, OMX, EarthRover, Reachy2, gamepads, keyboards, phones, OpenARM, **Unitree G1** (whole-body control added in v0.5.0) and the Seeed **reBot B601**.
    - Plugins are auto-discovered by package-name prefix (`lerobot_robot_*`, `lerobot_teleoperator_*`, `lerobot_camera_*`). Plugins exist for xArm, UR5e, Franka, AgileX Piper, WidowX, ARX5, I2RT YAM, GELLO, SpaceMouse, Quest, ROS 2 bridges, and tactile and depth cameras.
  - **Policies in the README:**
    - Imitation learning: ACT, Diffusion, VQ-BeT, Multitask DiT.
    - Reinforcement learning: HIL-SERL, TDMPC.
    - Vision-language-action models (VLAs): π0, π0-FAST, π0.5, **GR00T N1.7** (replaced N1.5 in v0.6), SmolVLA, X-VLA, EO-1, MolmoAct2, WALL-OSS, EVO1.
    - **World models:** VLA-JEPA, LingBot-VA, FastWAM, LaWAM, FLUX 3 Action.
    - **Reward models:** SARM, TOPReward, Robometer.
  - **LeRobotDataset v3.0** (lerobot ≥0.4.0):
    - Many episodes are packed per Parquet/MP4 file (v2 stored one file per episode).
    - Episode boundaries are resolved through Parquet metadata ("relational metadata").
    - `StreamingLeRobotDataset` streams directly from the Hub.
    - Optional **LanceDB** table backend.
  - **Data-relevant additions in v0.6.x:**
    - depth-map support, with depth units stored in the metadata;
    - separate RGB and depth codecs, plus libaom-AV1;
    - video re-encode and trim tools;
    - language columns replacing `subtask_index`;
    - a VLM subtask-annotation pipeline;
    - metadata-based episode filtering;
    - "record eval rollouts as LeRobot datasets";
    - DAgger smooth handover;
    - a 2× faster dataloader;
    - a `lerobot-rollout` command-line tool;
    - remote training on HF Jobs;
    - Foxglove visualization (Rerun was already supported);
    - an Isaac Teleop → SO-101 recording example.
  - **Benchmarks wired in through EnvHub:** LIBERO, MetaWorld, RoboCasa365, RoboTwin 2.0, RoboCerebra, RoboMME, LIBERO-plus and VLABench, with Docker smoke tests.
  - **v0.6 breaking changes:** the minimal `pip install` no longer includes dataset or training extras; the minimum PyTorch is 2.7; the RL stack was rebuilt (`sac` is now `gaussian_actor`).
  - **v0.5.0:** requires Python ≥3.12 and transformers v5.
- **Adoption / usage data:**
  - **77,473 HF datasets tagged LeRobot** (HF search page, today). Sep 2025 crawl: 16,065 public datasets (Kamenski). A TechTimes headline on 2026-05-25 said the Hub "surpasses 58,000 datasets in one year" (headline only).
  - June 2025 worldwide hackathon: **3,000+ participants, 44 countries, 250+ submissions** (LeRobot X post, headline only). A Munich node provided 50+ SO-101 arms. I could not verify a 2026 worldwide edition.
  - NVIDIA partnership: a 2026-07-18 article headline, "Nvidia and Hugging Face expand LeRobot…" (headline only). The release notes back this up with GR00T N1.7 integration, `nvidia/gr00t17-lerobot-libero_*` checkpoints and the Isaac Teleop example.
  - HF co-runs **RoboChallenge** with Dexmal, and Lightwheel shipped EgoSuite in partnership with HF.
- **Benchmarks / results (LeRobot docs):**
  - LIBERO: **π0.5 97.5% average** (97.0 / 99.0 / 98.0 / 96.0) versus OpenPI's own 96.85%.
  - GR00T N1.7: 96.5% average (preliminary, ≥50 episodes per suite).
  - LaWAM: 98.4 / 99.6 / 98.0 on spatial / object / goal.
- **Trade-offs vs alternatives:**
  - It is learning-centric and Python-first. It is not a real-time middleware: deterministic, low-latency control still needs ROS 2, dora or Copper.
  - It changes fast, with breaking renames even in v0.6.1 (`lerobot.types` → `lerobot.lerobot_types`).
  - Converting from v2.1 to v3.0 has corrupted episode boundaries in the wild (18.8% of successfully linted datasets in one audit, see Gaps).
  - Alternatives: openpi (Physical Intelligence models and recipes), and RLDS/TFDS (legacy, dormant).
- **Sources:**
  - via GitHits index: github.com/huggingface/lerobot (README.md, docs/source/lerobot-dataset-v3.mdx, libero.mdx, groot.mdx, robocasa.mdx, robocerebra.mdx, robotwin.mdx, and the v0.5.0/v0.6.0/v0.6.1 GitHub Release notes)
  - pkg_info pypi:lerobot
  - https://huggingface.co/datasets?other=LeRobot
  - https://www.kamenski.me/articles/analyzing-lerobot-datasets-on-hugging-face
  - https://thenextweb.com/news/uma-cadene-northstar-european-humanoid-robot
  - https://robotics.growbotics.ai/community/events/lerobot-hackathon-munich-2026

### ROS 2: the standard robot middleware and distribution
- **Status: ACTIVE.** **Lyrical Luth** reached general availability on **2026-05-22**. It is the 12th ROS 2 release, an LTS supported to **May 2031**, and its ROS Boss is Shane Loretz.
- **Who:** Open Robotics/OSRF and the community. RMW vendors include eProsima (Fast DDS, the default), ZettaScale (Zenoh), RTI, Eclipse Cyclone DDS and GurumDDS.
- **Links:** docs.ros.org; GitHub `ros2/ros2_documentation`.
- **Supported distros:**

  | Distro | Released | End of life |
  |---|---|---|
  | Lyrical | 2026-05-22 | May 2031 |
  | Kilted | 2025-05-23 | Dec 2026 |
  | Jazzy | 2024-05-23 | May 2029 |
  | Humble | 2022-05-23 | May 2027 |

- **License:** core packages are Apache-2.0 (not re-verified today).
- **What is new:**
  - Lyrical Tier-1 platforms are **Ubuntu 26.04 "Resolute" (amd64/aarch64)** and RHEL 10. Ubuntu Noble and Debian Trixie are Tier 3. The default RMW is still `rmw_fastrtps_cpp`.
  - Lyrical adds **`rosidl::Buffer` zero-copy publishing**: `uint8[]` fields become `rosidl::Buffer<uint8_t>` with pluggable backends (for example GPU/CUDA). It works with `rmw_fastrtps_cpp` and `rmw_zenoh_cpp`, and is the basis of Isaac ROS 5.0 dropping NITROS.
  - **rmw_zenoh_cpp** became Tier 1 and ships in the binaries from Kilted onward.
  - **MCAP** has been the default rosbag2 format since Iron (2023).
- **Adoption:** not re-measured today.
- **Trade-offs:** mature ecosystem with drivers, MoveIt and Nav2, but heavy for Python/ML workflows. Under rclpy, dora claims 10–17× lower latency (self-reported). The 5-year LTS cadence suits product companies.
- **Sources:** via GitHits github.com/ros2/ros2_documentation (source/Releases.rst, Releases/Release-Lyrical-Luth.rst, lyrical/release-timeline.rst, lyrical/supported-platforms.rst, Release-Kilted-Kaiju.rst, About-Different-Middleware-Vendors.rst, Release-Iron-Irwini.rst).

### MoveIt 2 / MoveIt Pro (PickNik): motion planning
- **Status: ACTIVE.**
  - **Qualcomm agreed to acquire PickNik** (announced **2026-09-23**; terms undisclosed). Qualcomm pledged to keep MoveIt 1 and 2 open source under their existing licenses, hardware-agnostic and with community roadmaps.
  - MoveIt Pro 9 shipped around April 2026, and 9.4.1 is dated 2026-07-07 (release-note URLs, headline only).
- **Who:** PickNik Robotics. Dave Coleman (CPO) is quoted; on the Qualcomm side, Nakul Duggal (EVP).
- **Links:** GitHub `moveit/moveit2`; docs.picknik.ai.
- **Stats:** about **2.0k stars, 785 forks**. Binary builds exist for Rolling, Lyrical, Jazzy and Humble, with stable branches for humble, jazzy and kilted.
- **License:** MoveIt 2 is BSD-3-Clause. MoveIt Pro is proprietary.
- **What it is:** an open-source planning, IK and collision stack. MoveIt Pro is a commercial runtime with behavior trees and perception-to-motion features (Google was the first customer).
- **Trade-offs:** the de facto open planner, but new ownership by a chip vendor is a strategic risk to watch (this is my inference, not reported). The GPU alternative is cuMotion (Isaac ROS).
- **Sources:** https://www.therobotreport.com/qualcomm-acquires-picknik-robotics-keep-moveit-open-source/ ; https://github.com/moveit/moveit2

### ros2_control: hardware abstraction and controller framework for ROS 2
- **Status: ACTIVE.** Branches exist for humble, jazzy and kilted; master serves Rolling and Lyrical.
- **Who:** ros-controls community (PickNik, Stogl Robotics and others).
- **Links:** GitHub `ros-controls/ros2_control`; control.ros.org.
- **Stats:** **935 stars, 455 forks, 232 tags.**
- **License:** Apache-2.0.
- **What it is:** a real-time controller manager plus hardware interfaces. It is the standard way to put a new arm or hand behind ROS 2.
- **Trade-offs:** the ROS-native choice. LeRobot uses its own Python robot classes instead, with optional ROS 2 bridges.
- **Sources:** https://github.com/ros-controls/ros2_control

### NVIDIA Isaac ROS: GPU-accelerated ROS 2 packages (perception, cuMotion)
- **Status: ACTIVE.**
  - **5.0.0 (2026-09-21):** migrated to Lyrical and **removed NITROS** in favor of native `rosidl::Buffer` with a CUDA backend.
  - 4.6.0 (2026-08-18): Jetson Orin support, Isaac Sim 6.0, and cloud-control workflows for the **Unitree G1**.
  - 4.5.0 (2026-07-06): sunset of the GXF implementation, a DNN stereo decoder, cuMotion 1.1.0.
  - 4.1 (2026-02-02, headline only).
- **Who:** NVIDIA.
- **Links:** GitHub `NVIDIA-ISAAC-ROS/isaac_ros_common`; nvidia-isaac-ros.github.io.
- **Stats:** isaac_ros_common has **319 stars, 227 forks**. Its last update was 2026-09-21 (GPU SM partitioning with CUDA MPS).
- **License:** Apache-2.0 for the packages (not re-verified). It depends on NVIDIA's CUDA/TensorRT stack.
- **Trade-offs:** the best Jetson Thor/Orin path, but tied to NVIDIA hardware and its frequent architecture churn (GXF, then NITROS, now Buffer).
- **Sources:** https://nvidia-isaac-ros.github.io/releases/index.html ; https://github.com/NVIDIA-ISAAC-ROS/isaac_ros_common

### Zenoh / rmw_zenoh: pub/sub/query protocol and ROS 2 middleware
- **Status: ACTIVE.** Zenoh 1.10.1 (2026-09-07), 1.10.0 (2026-08-14), 1.9.0 (2026-04-10).
- **Who:** Eclipse Foundation project, commercially backed by ZettaScale.
- **Links:** GitHub `eclipse-zenoh/zenoh`, `eclipse-zenoh/zenoh-python`, `ros2/rmw_zenoh`.
- **Stats:**
  - zenoh: **3.0k stars, 350 forks, 2.9M total crate downloads.**
  - zenoh-python: 173 stars.
  - rmw_zenoh: **497 stars, 111 forks.**
- **License:** EPL-2.0 OR Apache-2.0. rmw_zenoh is Apache-2.0.
- **What it is:** a brokerless or routed pub/sub layer that works well over WAN and Wi-Fi. dora uses it for shared memory and cross-machine links.
- **Trade-offs:** better than DDS on lossy, multi-site networks and for discovery storms. Fast DDS remains the default and the more battle-tested option.
- **Sources:** pkg_info crates:zenoh and pypi:eclipse-zenoh ; https://github.com/ros2/rmw_zenoh ; ROS docs as above.

### dora-rs: low-latency Rust dataflow framework for robotics and AI
- **Status: ACTIVE.** **1.0.0 shipped on 2026-09-02** and 1.0.1 on 2026-09-03.
  - The wire format and node APIs are frozen for all of 1.x.
  - 1.0 does **not** interoperate with 0.x.
- **Who:** dora-rs community (originated by Xavier Tao and Philipp Oppermann; from memory).
- **Links:** GitHub `dora-rs/dora`; blog docs/blog/2026-09-02-dora-1.0.md.
- **Stats:** **3.9k stars, 436 forks**; dora-cli has 58k total crate downloads.
- **License:** Apache-2.0 (crate). The PyPI node-API package lists MIT.
- **What it is:** a YAML-defined graph of nodes written in Rust, Python, C or C++.
  - Every message is an Apache Arrow array.
  - Messages ≥4 KB go over Zenoh shared memory, zero-copy.
  - There is a coordinator/daemon split, plus `record`/`replay`.
- **Benchmarks:** "node-to-node latency 10–17× lower than rclpy" on identical Python workloads (self-reported; examples/ros2-comparison).
- **Trade-offs:** much lighter than ROS 2 and ML-friendly through Arrow, but a smaller driver ecosystem.
- **Sources:** via GitHits github.com/dora-rs/dora docs/blog/2026-09-02-dora-1.0.md ; pkg_info crates:dora-cli, pypi:dora-rs.

### Copper-rs: deterministic Rust robotics runtime
- **Status: ACTIVE.** **1.0.0 (2026-07-02)**; 1.2.2 (2026-10-02).
- **Who:** Copper Robotics / copper-project.
- **Links:** GitHub `copper-project/copper-rs`.
- **Stats:** **1.5k stars, 106 forks**, 49k crate downloads.
- **License:** Apache-2.0.
- **What it is:** a compile-time scheduled task graph with structured logging and replay (the project describes itself as an "OS for physical AI").
- **Trade-offs:** strong determinism and replay for safety and debugging, but a young ecosystem.
- **Sources:** pkg_info and pkg_changelog crates:cu29.

### Rerun: multimodal, time-series logging, visualization and data platform
- **Status: ACTIVE.** 0.38.1 (2026-09-16), with about 10 releases between July and September 2026 and 133 versions in total.
- **Who:** Rerun (Stockholm). Seed of SEK 170M (about $17M) in March 2025, led by Point Nine. Customers include Meta, Google and Hugging Face (Dealroom/Techleap).
- **Links:** GitHub `rerun-io/rerun`; rerun.io.
- **Stats:** **11k stars, 849 forks.**
- **License:** MIT OR Apache-2.0.
- **What it is:** an SDK (Python, Rust, C++) plus a viewer for 3D, video and time-series data, with growing database features. LeRobot's visualizer depends on `rerun-sdk` (bumped to <0.34 in v0.6).
- **Trade-offs:** the best open viewer for learning datasets. Foxglove is stronger for ROS fleet observability.
- **Sources:** pkg_info pypi:rerun-sdk ; https://finder.techleap.nl/news/feed/rerun-raises-170m-for-smart-robots

### Foxglove / MCAP / Lichtblick: robotics observability, the log format, and the open fork
- **Status: ACTIVE.**
  - foxglove-sdk 0.28.0 (2026-09-29).
  - MCAP Python 1.5.0 (2026-09-24).
  - Lichtblick 1.29.1 (2026-09-08).
- **Who:**
  - Foxglove raised a **$40M Series B** (Nov 2025, headline only).
  - Lichtblick is led by **BMW AG** (its LICENSE says "Copyright 2024 BMW AG, 2021-2024 Foxglove").
- **Links:** GitHub `foxglove/mcap`, `foxglove/foxglove-sdk`, `lichtblick-suite/lichtblick`.
- **Stats:**
  - MCAP: **1.1k stars, 236 forks.**
  - foxglove-sdk: 311 stars.
  - Lichtblick: **1.1k stars, 751 forks**, 4.6k npm downloads/month.
- **License (status):**
  - **The Foxglove app is proprietary** and has been closed since 2024.
  - MCAP and foxglove-sdk are **MIT**.
  - Lichtblick is the **MPL-2.0** continuation of the open-source Foxglove Studio code ("open core").
- **What it is / adoption:**
  - MCAP is a self-describing, indexed container. It has been the **default rosbag2 format since ROS 2 Iron**.
  - EgoSuite-Open100K ships MCAP alongside LeRobot v3, and LeRobot 0.6 added Foxglove visualization.
- **Trade-offs:** MCAP is the raw-log lingua franca, while LeRobot v3 is the training format. A pipeline that converts MCAP to LeRobot is a natural product.
- **Sources:** pkg_info pypi:mcap, pypi:foxglove-sdk, npm:@lichtblick/suite ; via GitHits lichtblick README and LICENSE.

### Viser: web-based 3D visualization from Python
- **Status: ACTIVE.** 1.1.1 (2026-09-15), 1.1.0 (2026-08-16).
- **Who:** viser-project (Brent Yi et al.; originated in the nerfstudio ecosystem — from memory).
- **Stats:** **2.7k stars; 837k PyPI downloads/month.**
- **License:** MIT.
- **What it is / trade-offs:** an easy browser GUI for robot, scene and policy debugging. It is not a logging database (Rerun is).
- **Sources:** pkg_info pypi:viser.

### robot_descriptions.py: one-line import of about 100+ URDF/MJCF robot models
- **Status: ACTIVE.**
  - 3.2.0 (2026-09-12) added SRDFs and 5 new descriptions.
  - 3.1.0 (2026-07-23) added "GENE.01".
  - 3.0.0 (2026-07-11) added SRDF support.
- **Who:** Stéphane Caron and contributors (from memory).
- **Stats:** **833 stars, 75 forks.**
- **License:** Apache-2.0.
- **Sources:** pkg_info pypi:robot_descriptions.

### phosphobot: no-code control, recording and VLA-training app for low-cost arms
- **Status: SLOWING.** The last PyPI/GitHub release was 0.3.134 on **2025-10-22**. The repo is not archived.
- **Who:** phospho (Paris).
- **Stats:** **395 stars, 84 forks; 4.7k PyPI downloads/month.**
- **License:** MIT.
- **What it is:** supports SO-100/101, Koch, WX-250, Piper and Unitree Go2; teleoperation by keyboard, gamepad, leader arm or Quest; HF integration.
- **Trade-offs:** superseded in practice by LeRobot plus LeLab.
- **Sources:** https://github.com/phospho-app/phosphobot ; pkg_info pypi:phosphobot.

### YARP: middleware for iCub/IIT humanoids
- **Status: DORMANT for releases (unverified).** The releases page shows YARP 4.0.1 dated 2024-09-07 (the page summarizer has misread years before).
- **Who:** IIT robotology.
- **Stats:** **603 stars, 217 forks.**
- **License:** BSD-3-Clause, with optional LGPL/GPL components.
- **Trade-offs:** niche, tied to the iCub/ergoCub ecosystem.
- **Sources:** https://github.com/robotology/yarp ; https://github.com/robotology/yarp/releases

### Open-source data management, curation, annotation and capture tools
- **any4lerobot** (`Tavish9/any4lerobot`): **1.1k stars, MIT.**
  - Converts OXE, AgiBot World, RoboMIND, LIBERO and RoboCasa to LeRobot, and LeRobot to RLDS.
  - Migrates LeRobot versions, including v2.1 ↔ v3.0.
  - Likely ACTIVE, because v3.0 support postdates Oct 2025.
- **LeRobot built-ins:** delete, split, merge and add/remove features; recompute stats; re-encode and trim video; VLM subtask annotation; metadata filtering. LanceDB storage is documented as a backend.
- **LanceDB:** 11k stars, 7.9M PyPI downloads/month, Apache-2.0.
- **FiftyOne (Voxel51):** 11k stars, 108k downloads/month, Apache-2.0. Voxel51 publishes Egocentric-10K subsets in its dataset zoo.
- **RoboDM** (`berkeleyautomation/robodm`): 165 stars, v0.1.0 about a year ago. **DORMANT/SLOWING.**
- **RLDS:** 0.1.8, about 3 years old. **DORMANT** (legacy OXE stack).
- **trajlens:** v0.4.0, Apache-2.0. A new, tiny LeRobot dataset linter with 16 checks.
- **Capture hardware/software:**
  - **GELLO** (`wuphilipp/gello_software`): 538 stars, MIT. Supports I2RT YAM, Franka FR3/FER, UR and xArm, with gravity compensation. Its creator Philipp Wu is now CEO of XDOF.
  - **UMI** (`real-stanford/universal_manipulation_interface`): 1.6k stars, MIT. The core repo is largely static (likely DORMANT); the ecosystem lives on in forks.
- **Sources:** https://github.com/Tavish9/any4lerobot ; https://github.com/wuphilipp/gello_software ; https://github.com/real-stanford/universal_manipulation_interface ; pkg_info pypi:lancedb, fiftyone, robodm, rlds, trajlens.

---

# Part B — Open datasets

### LeRobot community datasets on the HF Hub: the long tail of crowdsourced robot data
- **Status: ACTIVE.** **77,473 datasets** match `other=LeRobot` today.
- **Profile (Kamenski crawl of 16,065 public datasets, Sep 2025):**
  - SO-100/101, Koch and LeKiwi make up **more than half**; about 296 robot-type strings appear.
  - 640×480 at 30 fps is used by about 72% of cameras.
  - Median **10 episodes** per dataset (mean 273); **about 43% have 1–5 episodes**.
  - Median episode length is 17.4 s.
  - **91% are single-task.**
  - The largest single dataset has 442,226 episodes.
- **Currently trending** (listing, downloads):
  - TacVerse/opendata: 6.1k (visuo-tactile)
  - EgoSteer/EgoSteer-RealWorld: 5.22k
  - physical-intelligence/libero: 43.3k
  - LightwheelAI/leisaac-pick-orange
- **License:** varies by uploader. Many datasets carry no license.
- **Trade-offs:** enormous breadth but shallow, with inconsistent calibration and quality (see Gaps).
- **Sources:** https://huggingface.co/datasets?other=LeRobot ; https://www.kamenski.me/articles/analyzing-lerobot-datasets-on-hugging-face

### Open X-Embodiment (OXE): pooled cross-embodiment robot data
- **Status: DORMANT** as a dataset (no new versions; RLDS/GCS). It is still widely used and has been converted to LeRobot.
- **Who:** Google DeepMind and 21 institutions (34 labs).
- **Links:** GitHub `google-deepmind/open_x_embodiment`; arXiv 2310.08864.
- **Size:** **1M+ real trajectories, 22 embodiments, 60 datasets, 527 skills (160,266 tasks).**
- **License:** code Apache-2.0; "all other materials" CC-BY 4.0. Constituent datasets keep their own terms.
- **Benchmarks:** RT-1-X +50% in the low-data regime; RT-2-X about 3× on emergent skills (paper).
- **Trade-offs:** breadth, but heterogeneous, with low-resolution, short and mostly-gripper data.
- **Sources:** https://robotics-transformer-x.github.io/ ; via GitHits OXE README.

### DROID: large in-the-wild Franka teleoperation dataset
- **Status: DORMANT.** Updates: language annotations in Dec 2024, improved calibrations for 36k episodes in Apr 2025. It remains the platform for RoboArena.
- **Who:** a multi-lab consortium (13 institutions, 50 collectors).
- **Links:** GitHub `droid-dataset/droid`; droid-dataset.github.io; HF port `lerobot/droid_1.0.1`.
- **Size:** **76k trajectories, 350 h, 564 scenes, 86 tasks.** 1.7 TB in RLDS; 8.7 TB raw stereo MP4.
- **Embodiment and modalities:** Franka Panda; 2× ZED 2 plus a wrist ZED Mini (stereo), with 1,417 viewpoints calibrated; Quest 2 teleoperation; 3 language annotations for 95% of successful episodes.
- **License:** the HF port is labeled Apache-2.0 (40M rows; 24 likes). The original release terms were not re-verified.
- **Trade-offs:** the best "diverse real scenes" single-arm set; single embodiment; no tactile or force data.
- **Sources:** https://droid-dataset.github.io/ ; https://huggingface.co/datasets/lerobot/droid_1.0.1 ; via GitHits docs/the-droid-dataset.md.

### BridgeData V2: WidowX tabletop data
- **Status: DORMANT** (2023).
- **Who:** UC Berkeley RAIL.
- **Size:** **60,096 trajectories** (50,365 teleoperated plus 9,731 scripted), 24 environments, 13 skills, 640×480 multi-view (including one RGB-D camera), with language.
- **License:** **CC BY 4.0** (commercial use OK).
- **Sources:** https://rail-berkeley.github.io/bridgedata/

### AgiBot World (Alpha/Beta): the largest open real-robot manipulation set
- **Status:** the dataset's last major release was Beta (2025-03-01). The ecosystem is ACTIVE (ICRA 2026 challenge, Genie Sim 3.0).
- **Who:** AgiBot and OpenDriveLab.
- **Links:** GitHub `OpenDriveLab/AgiBot-World`; HF `agibot-world/AgiBotWorld-Beta`, `agibot-world/AgiBotWorld-Alpha`; arXiv 2503.06669; GO-1 / GO-1 Air models.
- **Size:**
  - **Beta: 1,003,672 trajectories, 2,976.4 h,** about 43.8 TB (README) or 48.1 TB (HF card); 100 robots; 200+ task types; 87 atomic skills.
  - Alpha: 92,214 trajectories (8.5 TB).
- **Modalities:** multi-view RGB, depth, proprioception, force/torque. A subset has **visuo-tactile sensors and 6-DoF dexterous hands**; there are mobile dual-arm robots.
- **Format:** WebDataset, convertible to LeRobot.
- **License:** **CC BY-NC-SA 4.0 (non-commercial), gated.**
- **Adoption:** **94,323 downloads last month**, 80 likes.
- **Trade-offs:** unmatched scale, but non-commercial and single-vendor hardware.
- **Sources:** https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta ; via GitHits AgiBot-World README.

### RoboMIND: multi-embodiment teleoperation set from Beijing's humanoid innovation center
- **Status: SLOWING/unclear.** Current version is v1.2; "V2.0 announced on ModelScope" with no date verified.
- **Who:** X-Humanoid (Beijing Humanoid Robot Innovation Center) and collaborators. RSS 2025.
- **Links:** HF `x-humanoid-robomind/RoboMIND`; arXiv (Dec 2024).
- **Size:** **107k trajectories, 479 tasks, 96 object classes, 12.3 TB.**
  - Franka: 52,926
  - UR5e: 25,170
  - **Tien Kung humanoid: 19,152**
  - AgileX Cobot Magic: 10,629
- **Format:** HDF5.
- **License:** **Apache-2.0**, gated (accept conditions).
- **Adoption:** **53,240 downloads last month**, 54 likes.
- **Trade-offs:** commercially licensed and multi-embodiment, but with no depth in parts.
- **Sources:** https://huggingface.co/datasets/x-humanoid-robomind/RoboMIND

### Galaxea Open-World Dataset: mobile bimanual R1-Lite in homes, kitchens, retail and offices
- **Status: SLOWING/DORMANT.** Published 2025-08-30; no 2026 update verified.
- **Who:** Galaxea AI (G0 model; arXiv 2509.00576).
- **Size:** **500+ h, 227 tasks, 2.87 TB.**
- **Modalities:** 4 RGB streams (head, head-right, two wrists), joints, end-effector, IMU, chassis and torso actions; bilingual subtask language.
- **Format:** **LeRobot v2.1** (AV1, 15 fps).
- **License:** **CC BY-NC-SA 4.0, gated.**
- **Adoption:** **7,158 downloads last month**, 53 likes.
- **Sources:** https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset

### RH20T: contact-rich, multimodal real-robot dataset
- **Status: DORMANT** (ICRA 2024).
- **Who:** Shanghai Jiao Tong University.
- **Size:** **110k+ sequences, 147 tasks, 7 robot configurations, about 40 TB** (resized versions about 15 TB and 892 GB).
- **Modalities:** RGB-D, **6-DoF force/torque at 100 Hz, audio, fingertip tactile (configuration 7)**, IR, and paired human demonstration videos.
- **License:** **split.** Scenes 1–5 are CC BY-SA 4.0 (RH20T-C, commercial OK); scenes 6–10 are CC BY-NC 4.0 (RH20T-NC).
- **Trade-offs:** one of the few sets with force/torque, audio and tactile, but older robots and heavy to download.
- **Sources:** https://rh20t.github.io/

### Fourier ActionNet: humanoid upper-body teleoperation with dexterous hands
- **Status: SLOWING** (2025 release; no 2026 update found).
- **Who:** Fourier Intelligence and SJTU.
- **Links:** HF `FourierIntelligence/ActionNet`; action-net.org.
- **Size:** **30k+ trajectories, about 140 h.** Robots: GR1-T1, GR1-T2, GR2; **6-DoF and 12-DoF hands**; OAK-D cameras; egocentric (head) video plus state and action.
- **License:** **CC BY-NC-SA 4.0** for data, Apache-2.0 for code.
- **Sources:** https://action-net.org/

### Humanoid Everyday: diverse humanoid manipulation dataset
- **Status: SLOWING** (arXiv 2510.08807, Oct 2025; the cloud evaluation portal is "coming soon").
- **Who:** USC (Yue Wang) and Toyota Research Institute.
- **Size:** **10.3k trajectories, 3M+ frames, 260 tasks** at 30 Hz.
- **Modalities:** **RGB, depth, LiDAR, tactile** and language. Robot models were not listed on the page; I believe they are Unitree G1/H1, but this is unverified.
- **License:** not stated on the page.
- **Sources:** https://humanoideveryday.github.io/

### Unitree datasets on HF: G1 teleoperation datasets, updated continuously
- **Status: ACTIVE.** New G1 datasets were pushed "minutes ago" today.
- **Who:** Unitree Robotics.
- **Links:** HF org `unitreerobotics` (**162 datasets**). Models: UnifoLM-WMA-0, UnifoLM-ER-1/Flow (4B), UnifoLM-VLA.
- **What:** per-task G1 datasets with Dex1 grippers and **BrainCo or Inspire hands**, including **whole-body teleoperation ("WBT")** tasks. Head-stereo plus wrist cameras; LeRobot-style chunked MP4 and Parquet.
- **Example:** `G1_Dex1_HangCup` has 44 episodes, 8.8 GB and 1,089 downloads/month.
- **License:** not shown on the cards I opened.
- **Trade-offs:** small per-task sets (about 40–120 episodes), but the most active humanoid-specific open source.
- **Sources:** https://huggingface.co/unitreerobotics ; https://huggingface.co/datasets/unitreerobotics/G1_Dex1_HangCup

### NVIDIA Physical AI datasets (GR00T and others): sim, teleoperation and synthetic humanoid data
- **Status: ACTIVE.** 33 datasets match "PhysicalAI-Robotics".
- **Key items:**
  - **`nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim`:** about **325k sim trajectories** (240k GR1 humanoid tabletop, 72k arm kitchen, 9k bimanual Panda, plus others), 1.91 TB, **CC-BY-4.0**, **1,267,768 downloads last month**, 274 likes, 27 models trained on it.
  - **`nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL`** (2026-06-03, arXiv 2606.05160):
    - **22,312 physics-validated humanoid-object-interaction motions** for the **Unitree G1** (29 body DoF and 14 hand DoF), about 55 h, 224 GB.
    - Generated from synthetic video → SMPL-X/4D reconstruction → retargeting → RL validation in Isaac Lab.
    - **Apache-2.0** (bundled assets keep upstream licenses); 40,234 downloads last month.
  - **`nvidia/PhysicalAI-Robotics-Open-H-Embodiment`** (created Feb 2026):
    - surgical and ultrasound robotics, built with 30+ organizations;
    - 750 h, 120k trajectories, 4.54 TB, LeRobot v2.1, CC-BY-4.0;
    - 121,469 downloads last month.
- **Trade-offs:** commercially usable and very large, but mostly simulated or synthetic, so the sim-to-real gap applies.
- **Sources:** https://huggingface.co/datasets?search=PhysicalAI-Robotics&sort=downloads ; the three dataset pages above.

### BEHAVIOR dataset (BEHAVIOR-1K challenge demos): long-horizon household data in simulation
- **Status: ACTIVE.**
- **Who:** Stanford Vision and Learning Lab. Teleoperation via **JoyLo** whole-body interface; 2026 data collected by **Simovation**.
- **Links:** GitHub `StanfordVL/BEHAVIOR-1K`; HF `behavior-1k/2026-challenge-demos`.
- **Size:**
  - 2025: **10,000 demos, 1,200+ h, 50 tasks.**
  - 2026: **20,000 demos, 1,950 h, 100 tasks, 3.27 TB, LeRobot v3** (customized).
- **Modalities:** RGB-D, segmentation, object states, proprioception and actions, skill/subtask annotations. The robot is the Galaxea R1 (wheeled, bimanual) in OmniGibson.
- **License:** code MIT. The data/asset license was not verified.
- **Sources:** via GitHits BEHAVIOR-1K docs/challenge/index.md, archive/2025/index.md, dataset.md.

### EgoDex (Apple): dexterous egocentric video with 3D hand and upper-body poses
- **Status: SLOWING/DORMANT** (May 2025 paper).
- **Who:** Apple.
- **Links:** GitHub `apple/ml-egodex`; arXiv 2505.11709.
- **Size:** **829 h** of 30 Hz 1080p video from **Apple Vision Pro** (ARKit), **194 tabletop tasks**. Paired 3D head, upper-body and hand poses plus language. Splits: 725 h train, 7 h test, 97 h extra. Format: MP4 plus HDF5.
- **License:** **CC BY-NC-ND** (non-commercial, no derivatives).
- **Adoption:** used by H-RDT and OpenEgo; there is an HF viewer Space.
- **Trade-offs:** the best paired hand-pose quality at scale, but the license blocks commercial training.
- **Sources:** via GitHits apple/ml-egodex README.

### Ego4D / Ego-Exo4D (Meta and consortium): general egocentric video
- **Status: SLOWING.** Ego-Exo4D V2 and Ego4D v2.1 (Goal-Step) are the latest.
- **Size:**
  - Ego4D: about 3,670 h (paper; not re-verified today).
  - **Ego-Exo4D V2: 1,286.30 video hours (221.26 ego-hours), 5,035 takes**, Aria plus GoPro, with 3D.
- **License:** custom data license agreement (gated). The tooling code is MIT.
- **Trade-offs:** large and diverse, but not manipulation-focused, with few action labels.
- **Sources:** via GitHits github.com/facebookresearch/Ego4d README.

### HOT3D (Meta): egocentric 3D hand and object tracking
- **Status: DORMANT** (2024). It is still used in BOP and hand-tracking challenges.
- **Size:** **833 min, 3.7M+ images, 19 subjects, 33 objects**, from Aria and Quest 3. Optical mocap ground truth; MANO and UmeTrack hands; 6-DoF object poses.
- **License:** the full set is under the **HOT3D license agreement**. HOT3D-Clips are on HF (`bop-benchmark/hot3d`); OpenEgo lists them as CC BY-SA 4.0 (I did not check the HF card).
- **Sources:** https://facebookresearch.github.io/hot3d/ ; via GitHits hot3d README.

### EgoMimic, and its successor EgoVerse: human egocentric data co-trained with robot data
- **EgoMimic:**
  - **Status: DORMANT** (superseded).
  - GitHub `SimarKareer/EgoMimic`; small HF set `gatech/EgoMimic` (groceries, laundry, bowl tasks; Aria human plus robot HDF5).
- **EgoVerse** (arXiv 2604.07607; 2026-04-08, revised 2026-07-07):
  - **ACTIVE.**
  - **1,362 h, 80,000 episodes, 1,965 tasks, 240 scenes, 2,087 demonstrators** worldwide; standardized formats.
  - Lead author Ryan Punamiya plus 38 co-authors, academic and industry (I believe this is the Georgia Tech EgoMimic lineage, but that is unverified); egoverse.ai.
  - Finding: policies improve with more human data **only when it is aligned with the robot objective**.
  - Lightwheel's EgoSuite follows the EgoVerse consortium standards.
- **Sources:** https://arxiv.org/abs/2604.07607 ; via GitHits EgoMimic README.

### Build AI Egocentric-10K / Egocentric-100K: factory-worker egocentric video at massive scale
- **Status: SLOWING** for the public set (last updated 2026-02-16).
- **Who:** Build AI, which has raised about $15M (DreamVu; self-reported).
- **Links:** HF `builddotai/Egocentric-100K`, `builddotai/Egocentric-10K`, plus evaluation splits.
- **Size (100K):** **100,405 h; 2,010,759 clips; 10.8B frames**; 456×256 at 30 fps; monocular head-mounted fisheye. **Video only**: no audio, depth, IMU or hand pose.
- **License:** **Apache-2.0, gated** (contact information required).
- **Adoption:** **173,756 downloads last month**, 143 likes. The 10K set had 353 downloads.
- **Note:** DreamVu says Build AI reached "1M hours (Apr 2026, 14,228 workers)". The worker count matches the 100K set itself (100,405 h ÷ 7.06 h per worker ≈ 14.2k), so **treat the 1M-hour claim as unverified**.
- **Trade-offs:** volume leader with a commercial license, but low resolution and no action or pose labels.
- **Sources:** https://huggingface.co/builddotai ; https://huggingface.co/datasets/builddotai/Egocentric-100K ; https://www.dreamvu.ai/blog/robot-training-data-companies-2026

### Lightwheel EgoSuite-Open100K: egocentric human data with hand and body pose
- **Status: ACTIVE.** Released **2026-08-26**.
- **Who:** Lightwheel (Chinese name 光轮 / Guanglun Intelligence), in partnership with HF.
- **Size:** **10,000 h live; 100,000 h targeted.** 15,000+ tasks and 15,000+ scenes across 7 categories and 128 scene types. Head-mounted cameras, plus a wrist camera in the "EgoPro" variant. **Hand pose**, full-body pose in some subsets, event-level semantics.
- **Format:** **LeRobot v3 (streamable) plus MCAP.**
- **License:** "academic research and **commercial training**" per the blog; per-card terms apply. One card surfaced in search as "Request access to EgoSuite-Open100K" (`LightwheelAI/EgoStandard`, not opened), which suggests gating.
- **Sources:** https://huggingface.co/blog/LightwheelAI/egosuite-open100k

### Open-AoE (Ant Group): smartphone-captured egocentric manipulation data and toolchain
- **Status: ACTIVE** (July 2026; arXiv 2607.14183).
- **Size:** **about 2,000 h, 8,000+ atomic-action tasks, 500+ contributors, 400+ phone models.**
- **Modalities:** RGB, **MANO 21-joint bimanual hand pose**, 6-DoF camera trajectories, English action annotations, scene and object text.
- **License:** **CC BY 4.0.** Hosted on GitHub, HF and ModelScope.
- **Trade-offs:** cheap, scalable consumer capture with a commercial license, but no depth or tactile data and pose that is estimated, not measured.
- **Sources:** https://arxiv.org/html/2607.14183

### HA-Ego-500 (Human Archive): densely annotated workplace egocentric data
- **Status: ACTIVE** (Aug 2026).
- **Size:** **500 h** (519.3 h collected), **310k+ labeled steps**, 60+ work environments; median step about 4 s; hands engaged in 92% of steps. Custom multi-camera rigs.
- **License:** not disclosed.
- **Sources:** https://ego500.humanarchive.ai/

### OpenEgo: a unified egocentric-manipulation corpus
- **Status: SLOWING** (Sep 2025).
- **Who:** UT Dallas and Physical Automation.
- **Size:** **1,107 h, 290 tasks, 600+ environments**, merging 6 sources (CaptainCook4D, HOI4D, HoloAssist, EgoDex, HOT3D, HO-Cap) with standardized MANO-21 hands.
- **License:** inherits each source's license (mixed, including NC and ND).
- **Sources:** https://arxiv.org/html/2509.05513

### AMASS: unified human mocap in SMPL
- **Status: DORMANT** (ICCV 2019). Still foundational for humanoid retargeting.
- **Size:** **40+ h, 300+ subjects, 11k+ motions, 15 mocap datasets.**
- **License:** registration plus the MPI license. It is a research license; commercial use requires separate terms (from memory; verify).
- **Sources:** https://amass.is.tue.mpg.de/

### LAFAN1 (Ubisoft): game-quality mocap
- **Status: DORMANT.** Shot in 2017.
- **Size:** **5 subjects, 77 sequences, 496,672 frames at 30 fps (about 4.6 h)**, BVH.
- **License:** **CC BY-NC-ND 4.0.**
- **Note:** widely retargeted to G1/H1 for locomotion RL. Retargeted derivatives inherit the non-commercial, no-derivatives terms (my reading of the license).
- **Sources:** via GitHits ubisoft-laforge-animation-dataset README and license.txt.

### OMOMO: full-body human-object interaction mocap
- **Status: DORMANT** (2023).
- **Size:** about 10 h (paper; not re-verified), 15 large objects (10 train, 5 test).
- **License:** code MIT; it requires SMPL-H/SMPL-X, which are under non-commercial model licenses.
- **Sources:** via GitHits lijiaman/omomo_release.

### Motion-X / Motion-X++: whole-body (SMPL-X) motion with text
- **Status: DORMANT.** Motion-X++ was reorganized on HF in Mar 2025.
- **License:** **non-commercial research only**; access is by form. Original RGB videos are not redistributed.
- **Sources:** via GitHits IDEA-Research/Motion-X README and LICENSE.

### New 2026 robot-action datasets
- **Axis Sim Dataset V1** (Axis Robotics, **2026-09-04**):
  - **50k+ human-teleoperated sim trajectories, 60k+ scene variants, 207 tasks**, Franka FR3.
  - Crowdsourced through a browser platform, "Axis Hub", that claims 200k+ contributors and describes itself as a "top-3 dApp on Base".
  - Claims **160k+ downloads** (self-reported). HF `axisrobotics/Franka-Dataset`. The license is called "open-source" without specifics.
  - $12M seed (Hack VC).
  - Source: https://www.globenewswire.com/news-release/2026/09/04/3356410/0/en/axis-robotics-open-sources-one-of-the-largest-franka-arm-simulation-datasets-for-physical-ai.html
- **ABC / "ABC-130K"** (XDOF with UC Berkeley BAIR, June 2026):
  - **130k manipulation trajectories, 300 h simulation, 100 h evaluations** (TechCrunch).
  - DreamVu calls it the "largest open-source bimanual manipulation dataset".
  - HF ID and license not verified.
- **MolmoAct 2 data** (AI2, 2026-05-05):
  - **720+ h bimanual YAM demonstrations**, plus all training datasets and evaluation rollouts released; language labels grew from about 71k to 146k.
  - Results: LIBERO 97.2% (98.1% with depth reasoning); Franka real-world 87.1%; third-party **Cortex AI** evaluation score 0.51, first on 7 of 8 tasks.
  - License not verified.
  - Source: https://allenai.org/blog/molmoact2
- **TacVerse/opendata** (Xense Robotics):
  - **Visuo-tactile bimanual** manipulation in LeRobot format; about 40M rows and about 1.28k episodes; **CC-BY-SA-4.0**; about 6.1k downloads.
  - Source: https://huggingface.co/datasets/TacVerse/opendata
- **RoboArena data dump** (2026-02-03):
  - **23,629 evaluation-session rows** with rollouts, success scores and preference labels, plus 3-camera video and proprioception; 18.5 GB; **MIT**; 1,018 downloads last month.
  - This is a rare **failure-inclusive** real-robot corpus.
  - Source: https://huggingface.co/datasets/RoboArena/DataDump_02-03-2026
- **RoboCerebra unified:** `lerobot/robocerebra_unified`, LeRobot v3; 6,660 episodes, 571,116 frames, 1,728 language subtasks (long-horizon).

---

# Part C — Benchmarks and competitions

### BEHAVIOR Challenge (Stanford): long-horizon household tasks in OmniGibson
- **Status: ACTIVE.** The 2026 edition launched 2026-07-02. **Submission deadline 2026-10-16; winners announced 2026-11-04.**
- **What is measured:**
  - 2026: 100 full household tasks in 7 scenes (4 new). One track with RGB, depth and proprioception. Ranked by average task success with **BDDL partial credit ("q-score")**. Baselines are π0.5 and GR00T N1.7. Prize pool **$11k** ($5k / $3k / $2k / $1k open source).
  - 2025: NeurIPS 2025, 50 tasks, **18 teams** from 4 countries.
- **2025 results:**
  - **1st, "Robot Learning Collective"** (independent researchers Ilia Larchenko, Gleb Zarin, Akash Karnatak): **q-score 26%** on public and private test sets; binary success 11.2% (public) and 12.4% (private). Built on π0.5 with correlated-noise flow matching and stage tracking. Compute: about $13k (8×H200 for about 15 days plus fine-tuning).
  - **2nd, Team Comet (openpi-comet): q-score 0.2514** (held-out). Post-challenge they report 0.345 on public validation (self-reported).
- **Sources:**
  - via GitHits StanfordVL/BEHAVIOR-1K docs/challenge/index.md and archive/2025/index.md
  - https://arxiv.org/html/2512.06951v2
  - via GitHits mli0603/openpi-comet README

### RoboArena: distributed, double-blind, pairwise real-robot evaluation on DROID
- **Status: ACTIVE.** Latest public data dump 2026-02-03; the leaderboard site is live but JavaScript-rendered, so I could not read it.
- **Who:** 27 authors across 7 academic institutions (arXiv 2506.18123, June 2025).
- **What is measured:** crowd-sourced A/B preferences and progress scores of generalist policies on evaluator-chosen tasks, served remotely through a `BasePolicy` server.
- **Results:** the paper used 600+ pairwise episodes across 7 policies. The current leaders were not verified (site unreadable).
- **Trade-offs:** scalable and hard to game, but tied to a single platform (Franka/DROID).
- **Sources:** https://arxiv.org/abs/2506.18123 ; https://huggingface.co/datasets/RoboArena/DataDump_02-03-2026 ; via GitHits robo-arena/roboarena README.

### RoboChallenge (Dexmal and Hugging Face): large-scale real-robot evaluation as a service
- **Status: ACTIVE.** Table30 V2 launched 2026-03-24; there is a CVPR 2026 workshop paper.
- **What is measured:**
  - Fleet of **10 robots of 4 types** (UR5, Franka, Cobot Magic Aloha, ARX-5).
  - Table30 has 30 tasks. Submission is by remote API: the model runs on the participant's side.
  - Up to 1,000 fine-tuning episodes per task.
  - Metrics: success rate plus progress score (0–10 per task over 10 rollouts).
  - V2 adds zero-shot and out-of-domain tests and a DOS-W1 mobile platform.
- **Results (snapshots, not directly comparable across versions):**
  - Initial (Oct 2025): **π0.5 43.7% success / 62.2 score**; π0 28.3%; CogACT 11.7%.
  - **Spirit v1.5 (Spirit AI) ranked #1** in Jan 2026, and its weights and code were then open-sourced.
  - Aggregator snapshot (Feb 2026): **DM0 (Dexmal) 37.3% success / 49.08**; π0.5 17.67%; π0 9%.
  - Table30 V2 (Mar 2026, Chinese press): **DM0 62%**, GigaBrain-0.1 about 52%, π0.5 42.67%, RDT-1B 15%. Multi-step tasks such as sandwich making are near 0%.
  - **Conflict of interest:** the organizer's own model (DM0) leads.
- **Sources:**
  - https://arxiv.org/html/2510.17950v1
  - https://finance.yahoo.com/news/robochallenges-top-ranked-embodied-ai-064100221.html
  - https://www.sota2.com/research/sota/overall-robotic-manipulation-on-table30-robochallenge
  - https://www.leaderobot.com/news/7455

### AgiBot World Challenge: manipulation and world-model competition
- **Status: ACTIVE.** ICRA 2026 finals were held in **Vienna on 2026-06-05**; there was an IROS 2025 edition before.
- **What is measured:**
  - **Reasoning-to-Action (R2A)** track: understanding, planning and execution.
  - **World Model (WM)** track: predicting physical change.
  - Online automated rounds (Genie Sim 3.0, EWMBench), then **real-robot finals on the AgiBot G2** humanoid.
- **Results:**
  - **526 teams from 27 countries.**
  - **R2A:** 1st **PrismBot (vivo)**, 2nd RP-VLA (Shanghai RoboParty), 3rd GreenVLA.
  - **WM:** 1st **NeoVerse-ABot** (CAS / Amap CV Lab), 2nd PAI@IAII, 3rd Loop (USTC).
  - Scores were not disclosed.
- **Sources:** https://www.prnewswire.com/news-releases/agibot-world-challenge-2026-advances-embodied-ai-competition-from-simulation-to-real-robot-testing-at-icra-2026-302792634.html

### RoboTwin 2.0: bimanual simulation benchmark and data generator
- **Status: ACTIVE.**
  - ICML 2026.
  - RMBench (memory-dependent tasks) on 2026-03-03.
  - StarVLA support on 2026-02-20.
  - IsaacLab-Arena and RLinf support on 2026-01-23.
  - Integrated in LeRobot v0.6.
- **Who:** RoboTwin-Platform team (Tianxing Chen et al.).
- **Links:** GitHub `RoboTwin-Platform/RoboTwin`; arXiv 2506.18088; HF `TianxingChen/RoboTwin2.0`.
- **What is measured:** 50 tasks on Aloha-AgileX. Train on 50 clean demos per task, evaluate 100 episodes per task. **Easy (clean) vs Hard (domain-randomized)** settings. 100k+ pre-collected trajectories, 731 objects, 5 embodiments.
- **Results (paper baselines):**

  | Policy | Easy | Hard |
  |---|---|---|
  | π0 | 46.4% | 16.3% |
  | RDT | 34.5% | 13.7% |
  | DP3 | 55.2% | 5.0% |
  | ACT | 29.7% | 1.7% |
  | DP | 28.0% | 0.6% |

  The live leaderboard was unreadable (JavaScript).
- **Sources:** https://arxiv.org/html/2506.18088 ; via GitHits RoboTwin README ; https://robotwin-platform.github.io/leaderboard (protocol only).

### ManiSkill3: GPU-parallel manipulation simulation and benchmark
- **Status: ACTIVE.** mani-skill 3.0.1 (2026-04-21); repo pushed 2026-08-04.
- **Who:** Hao Su lab (UCSD) / Hillbot. RSS 2025 paper.
- **Stats:** **3.3k stars, 550 forks; 24k PyPI downloads/month.**
- **License:** Apache-2.0.
- **Results:** 200k+ state-only FPS and 30k+ FPS with rendering on a single RTX 4090 (self-reported).
- **Sources:** pkg_info pypi:mani_skill ; via GitHits ManiSkill docs.

### LIBERO (brief): lifelong-learning simulation suites, now saturated
- **Status:** the original repo is effectively DORMANT but remains the default sanity check.
- **Stats:** **2.3k stars, 476 forks.** 130 tasks in 4 suites. Code MIT, data CC BY 4.0.
- **Results:** π0.5 96.85% (OpenPI) and 97.5% (LeRobot); GR00T N1.7 96.5%; MolmoAct 2 97.2–98.1%.
- **Successors:** LIBERO-plus (robustness) and RoboCerebra.
- **Sources:** https://github.com/Lifelong-Robot-Learning/LIBERO ; LeRobot docs via GitHits.

### HumanoidBench: simulated humanoid locomotion and manipulation
- **Status: likely DORMANT.** 45 commits total, no releases; the last commit date could not be verified.
- **Who:** UC Berkeley (Sferrazza et al., RSS 2024).
- **Stats:** **799 stars, 130 forks**, MIT.
- **What is measured:** 15 manipulation and 12 locomotion tasks (the page also cites 31 including variants) on a Unitree H1 with Shadow hands (G1 and Digit variants), using MuJoCo/MJX. Baselines: TD-MPC2, DreamerV3, SAC, PPO.
- **Sources:** https://github.com/carlosferrazza/humanoid-bench

### World Humanoid Robot Games: humanoid athletics and skills competition
- **Status: ACTIVE.** The 2nd edition ran **2026-08-22 to 08-26** in Beijing.
- **Results (as reported by the Beijing municipal government; event rules not verified):**
  - **666 teams, 16 countries, 2,000+ robots** — +138% teams and 4× robots versus 2025.
  - **100 m in 9.39 s** (Tianzhuo team; 21.5 s in 2025).
  - High jump 2.8843 m (Tiangong).
  - Long jump over 7 m.
  - Football moved from 3v3 to 5v5.
- **Sources:** https://english.beijing.gov.cn/beijinginfo/sci/latesttrends/202608/t20260825_4836357.html

### RoboCup Humanoid League: robot soccer
- **Status: ACTIVE.** RoboCup 2026 was held in **Incheon in July 2026**. It was the first season after the Humanoid League and the Standard Platform League merged.
- **Winners:** Small: **Invic (Wuhan University)**; Middle: **B-Human (University of Bremen and DFKI)**; Large: **Tsinghua Hephaestus**.
- **Sources:** https://letsdatascience.com/news/robocup-2026-humanoid-league-declares-division-winners-ffe56ef5

### New 2026 benchmarks and evaluation efforts
- **ArmnetBench v0.1** (Armnet, arXiv 2607.24481, July 2026):
  - 4 SO-101 arms in 3 cells (**$359 single-arm cell, $477 bimanual cell**); 12 tasks; 7 policies; 2,518 rollouts.
  - Leader: **π0.5 at 47.6% strict success** (45.4% single-arm, 52.1% bimanual).
  - Source: https://arxiv.org/html/2607.24481v1
- **RoboCasa365:** 365 kitchen tasks (about 65 atomic, about 300 composite).
- **RoboCerebra:** long-horizon.
- **RoboMME, LIBERO-plus, VLABench:** integrated into LeRobot v0.6.
- **RMBench:** memory-dependent tasks, on top of RoboTwin.
- **Cortex AI:** third-party real-robot evaluation, used by AI2 for MolmoAct 2.
- **Not opened (budget ran out):** RoboWorld (neural-simulator policy evaluation, arXiv 2607.01060) and RobotArena∞ (real-to-sim benchmarking).

---

# Part D — The data economy

### Who supplies data (2026)
Unless marked otherwise, these entries are from DreamVu's 2026-08-06 landscape post. DreamVu is itself a vendor, so treat them as self-reported.

| Company | Data / product | Funding and scale (as reported) |
|---|---|---|
| **XDOF** | Teleoperation pipelines on deployment robots, GELLO-style capture, planned egocentric wearables; ABC dataset | **$70M** (Thrive, Spark, a16z, Lux, WndrCo); about 60 staff; 20 customers incl. "several frontier labs" (TechCrunch, 2026-06-17); founder Philipp Wu (GELLO) |
| **Scale AI** | Teleoperation and human demos | Customers: Physical Intelligence, Generalist, Cobot |
| **Lightwheel** | SimReady OpenUSD assets, EgoSuite, RoboFinals | Customers named: Google DeepMind, Figure, AgiBot, ByteDance, Geely, BYD |
| **Mecka AI** | Body-worn sensors and iPhones; full-body kinematics | About $68M total; claims a $100M run-rate in signed contracts; 1X is a customer (Flikforge says $60M in June 2026) |
| **Config** | Bimanual capture plus model | $27M seed at a $200M valuation (Samsung ecosystem); targets 1M hours |
| **Human Archive** | Head rigs with RGB-D, tactile gloves, mocap suits | $8.2M (Wing, NVP, YC W26); 1,000+ active headsets; **pays workers $1/h base, while competitors pay ₹250–400/h ($2.63–4.20)** (TechCrunch, 2026-05-26) |
| **Build AI** | Factory egocentric video | About $15M; open Egocentric-10K/100K |
| **Encord** | Curation and annotation tooling (LiDAR, video) | $110M total, including a $60M Series C (Feb 2026); physical-AI revenue up 10× |
| **Bones Studio** | Studio optical mocap | BONES-SEED, 142k+ sequences (used for NVIDIA SONIC) |
| **Luel** | Rights-cleared marketplace (incl. curated Ego4D/Ego-Exo4D) plus custom collection | Lightspeed-backed |
| **Axis Robotics** | Crowdsourced browser teleoperation in simulation | $12M seed (Hack VC) |
| **Genesis AI** | **Tactile e-skin data glove** ("100× cheaper hardware, up to 5× collection efficiency" vs teleoperation; self-reported) plus egocentric and internet video plus sim | $105M seed; nothing open-sourced (2026-05-06) |
| **MANUS** | Data gloves (Metagloves Pro Haptic) | Users include ROBOTERA, BrainCo, TESOLLO, Xynova, Shadow Robot; retargeting for AgiLink, Inspire, Allegro (ICRA 2026 blog) |
| Also (headline only) | Midcentury | $15M seed |
| Also (headline only) | vision lab (factory data) | $6M seed |
| Also (headline only) | Proception | $11M seed |
| Also (headline only) | Tacta Systems | "large-scale skill capture" |

**Closed-data benchmark for scale:** Generalist's GEN-0 was trained on **270,000 h of real-world manipulation data, growing by more than 10,000 h per week**, collected through "data foundry partners" (Generalist blog, 2025-11-04). That is 2–3 orders of magnitude more than the largest open real-robot sets.

### Price points (all vendor-published; units are not comparable — raw vs accepted vs training-ready hours)

| Source (vendor, date) | Item | Price |
|---|---|---|
| DexSet (via EXYLOS, 2026-09-05) | Egocentric video | $15–40 per raw hour |
| DexSet | Teleoperation | $28–60 per raw hour |
| DexSet | Annotation | +$8–25 per data hour |
| Robotics Center of Silicon Valley | Production dataset with 500+ demos | $50k–200k |
| Robotics Center of Silicon Valley | Per episode | $8–35 |
| DataX Power (2026-07-07) | Single-arm teleoperation | $15–30/h |
| DataX Power | Bimanual ALOHA-style | $40–80/h |
| DataX Power | Egocentric wearable | $25–60/h |
| DataX Power | **Full humanoid multi-sensor** | **$80–150/h**, or $50–150 per complex demo |
| DataX Power | QA overhead | 75% acceptance → 1.33× cost |
| DreamVu | Robot-specific teleoperation | $50–200/h; Build AI's open data is "effectively free" |
| Flikforge (CEO op-ed, 2026-07-07) | Raw footage resale | **$2–5/h**, versus $5–20/h to create; buyers discard about 90% |

**Formats that dominate:**
- **LeRobotDataset** (v2.1 moving to v3.0) for training and Hub distribution.
- **MCAP** for raw ROS/robot logs.
- Legacy and other formats: RLDS/TFDS (OXE), HDF5 (RoboMIND, EgoDex, ALOHA-style), WebDataset (AgiBot), VRS (Meta Aria), Zarr (UMI/Diffusion Policy).

---

## Also notable (brief)

- **LeLab** (`huggingface/leLab`), web UI for LeRobot — ACTIVE (added to the README in 2026).
- **Spirit v1.5** (Spirit AI), open VLA that ranked #1 on RoboChallenge in Jan 2026 — ACTIVE.
- **GR00T N1.7 checkpoints** for LeRobot LIBERO (`nvidia/gr00t17-lerobot-libero_*`) — ACTIVE.
- **StarVLA**, **RLinf**, **XPolicyLab/RoboDojo**: VLA/RL infrastructure that integrates RoboTwin and BEHAVIOR — ACTIVE (2026 updates).
- **Genie Sim 3.0 / EWMBench** (AgiBot): simulation and world-model evaluation used in the ICRA 2026 challenge — ACTIVE.
- **Open-H-Embodiment**: surgical/ultrasound LeRobot data from 30+ organizations — ACTIVE.
- **EgoSteer/EgoSteer-RealWorld**: trending LeRobot-tagged dataset (5.22k downloads); contents not verified.
- **Reachy 2, HopeJR, SO-101, LeKiwi, OMX**: HF-supported open hardware — ACTIVE.
- **UnifoLM-WMA/ER/VLA**: Unitree's open models — ACTIVE.

---

## Comparison tables

### (a) Tooling

| Project | Role | License | Stars | Latest release | Status |
|---|---|---|---|---|---|
| LeRobot | Learning library, dataset format, Hub | Apache-2.0 | 27k | v0.6.1 (2026-08-03) | ACTIVE |
| ROS 2 | Middleware/distro | Apache-2.0 (core) | — | Lyrical (2026-05-22) | ACTIVE |
| MoveIt 2 | Motion planning | BSD-3 | \~2.0k | binaries for Lyrical | ACTIVE (Qualcomm acquiring PickNik) |
| ros2_control | Controllers/HAL | Apache-2.0 | 935 | 232 tags | ACTIVE |
| Isaac ROS | GPU ROS packages | Apache-2.0* | 319 (common) | 5.0.0 (2026-09-21) | ACTIVE |
| Zenoh / rmw_zenoh | Transport/RMW | EPL-2.0/Apache-2.0 | 3.0k / 497 | 1.10.1 (2026-09-07) | ACTIVE |
| dora-rs | Dataflow runtime | Apache-2.0 | 3.9k | 1.0.1 (2026-09-03) | ACTIVE |
| Copper-rs | Deterministic Rust runtime | Apache-2.0 | 1.5k | 1.2.2 (2026-10-02) | ACTIVE |
| Rerun | Logging/visualization | MIT/Apache-2.0 | 11k | 0.38.1 (2026-09-16) | ACTIVE |
| MCAP | Log format | MIT | 1.1k | py 1.5.0 (2026-09-24) | ACTIVE |
| Foxglove app | Observability | Proprietary | — | — | ACTIVE (closed) |
| Lichtblick | Open Foxglove fork | MPL-2.0 | 1.1k | 1.29.1 (2026-09-08) | ACTIVE |
| Viser | 3D web visualization | MIT | 2.7k | 1.1.1 (2026-09-15) | ACTIVE |
| robot_descriptions | Robot models | Apache-2.0 | 833 | 3.2.0 (2026-09-12) | ACTIVE |
| phosphobot | Low-cost arm application | MIT | 395 | 0.3.134 (2025-10-22) | SLOWING |
| YARP | iCub middleware | BSD-3 | 603 | 4.0.1 (2024?, unverified) | DORMANT? |
| any4lerobot | Format converters | MIT | 1.1k | — | likely ACTIVE |
| RLDS | Legacy dataset library | Apache-2.0 | — | 0.1.8 (about 3 years ago) | DORMANT |

\*Not re-verified.

### (b) Datasets

| Dataset | Size | Embodiment / capture | Modalities | License (commercial use?) | Year / last update | HF downloads per month |
|---|---|---|---|---|---|---|
| AgiBot World Beta | 1.0M trajectories / 2,976 h | AgiBot fleet (100 robots), dexterous subset | RGB-D, force/torque, visuo-tactile subset | CC BY-NC-SA (no) | 2025 | 94,323 |
| OXE | 1M+ trajectories | 22 embodiments | RGB, varied | CC-BY 4.0 (yes; check sub-datasets) | 2023 | — |
| GR00T X-Emb Sim | \~325k trajectories | GR1, Panda, G1 (sim) | Sim RGB/state | CC-BY-4.0 (yes) | 2025 | 1,267,768 |
| RH20T | 110k+ sequences | 7 arm configurations | RGB-D, force/torque, audio, tactile | split: BY-SA / BY-NC | 2023–24 | — |
| RoboMIND | 107k trajectories | Franka, UR5e, Tien Kung, AgileX | RGB, state | Apache-2.0 (yes; gated) | v1.2 | 53,240 |
| DROID | 76k trajectories / 350 h | Franka | Stereo RGB, language | HF port Apache-2.0 | 2024–25 | — |
| BridgeData V2 | 60,096 trajectories | WidowX | RGB-D, language | CC BY 4.0 (yes) | 2023 | — |
| Axis Sim V1 | 50k+ trajectories | FR3 (sim) | Sim | "open" (unclear) | 2026-09 | 160k+ (self-reported) |
| Fourier ActionNet | 30k+ trajectories / 140 h | GR1/GR2 + 6/12-DoF hands | Egocentric RGB, state | BY-NC-SA (no) | 2025 | — |
| BEHAVIOR 2026 | 20k demos / 1,950 h | R1 (sim) | RGB-D, segmentation, subtasks | not verified | 2026 | — |
| Humanoid Everyday | 10.3k trajectories / 260 tasks | Humanoid | RGB-D, LiDAR, tactile | not stated | 2025 | — |
| Galaxea | 500+ h / 227 tasks | R1-Lite | 4×RGB, IMU, language | BY-NC-SA (no) | 2025 | 7,158 |
| GRAIL | 22,312 motions / \~55 h | G1 (sim) | SMPL-X human-object interaction + G1 | Apache-2.0 (yes) | 2026-06 | 40,234 |
| Unitree HF | 162 datasets | G1 + Dex1/BrainCo/Inspire | Stereo + wrist | not shown | 2026 (daily) | \~1k per set |
| MolmoAct 2 YAM | 720+ h | Bimanual YAM | RGB, language | not verified | 2026-05 | — |
| TacVerse | \~1.28k episodes | Bimanual | **Visuo-tactile** | CC-BY-SA (yes, share-alike) | 2026 | \~6.1k |
| Egocentric-100K | 100,405 h | Human head-mounted fisheye | Video only, 456×256 | Apache-2.0 (yes; gated) | 2025 / Feb 2026 | 173,756 |
| EgoSuite-Open100K | 10k h (→100k) | Human head (+wrist) | Hand and body pose | commercial training allowed | 2026-08 | — |
| Open-AoE | \~2,000 h | Smartphones | MANO hands, camera pose | CC BY 4.0 (yes) | 2026-07 | — |
| EgoVerse | 1,362 h / 80k episodes | Human egocentric | Annotations | see egoverse.ai | 2026 | — |
| Ego-Exo4D V2 | 1,286 h | Aria + GoPro | Multi-view, 3D | custom agreement | 2024–25 | — |
| OpenEgo | 1,107 h | Mixed (6 sources) | MANO-21 | mixed | 2025 | — |
| EgoDex | 829 h | Vision Pro | 3D hand/body pose | BY-NC-ND (no) | 2025 | — |
| HA-Ego-500 | 500 h | Multi-camera rig | Dense step labels | not disclosed | 2026-08 | — |
| HOT3D | 13.9 h (833 min) | Aria / Quest 3 | Mocap ground truth for hands and objects | HOT3D agreement | 2024 | — |
| AMASS | 40+ h | Optical mocap | SMPL | research license | 2019 | — |
| LAFAN1 | 4.6 h | Mocap | BVH | BY-NC-ND (no) | 2017/2020 | — |

### (c) Benchmarks

| Benchmark | What is measured | Operator | Current leader and score (date) | Status |
|---|---|---|---|---|
| BEHAVIOR | 50 → 100 long-horizon household tasks (sim), q-score | Stanford SVL | 2025: Robot Learning Collective, q 26%; 2026 results due 2026-11-04 | ACTIVE |
| RoboChallenge Table30 (V2) | Real-robot fleet, success rate + progress | Dexmal + HF | DM0 \~62% on V2 (Mar 2026, press); Spirit v1.5 #1 (Jan 2026) | ACTIVE |
| RoboArena | Pairwise real-robot preferences (DROID) | 7-institution consortium | Not verified (site JavaScript) | ACTIVE |
| AgiBot World Challenge | R2A + world model; sim → G2 robot finals | AgiBot | PrismBot (vivo) R2A; NeoVerse-ABot WM (Jun 2026) | ACTIVE |
| RoboTwin 2.0 | 50 bimanual tasks, clean vs randomized | RoboTwin team | Paper: π0 46.4% / 16.3% | ACTIVE |
| ArmnetBench | 12 tasks on SO-101 arm farm | Armnet | π0.5 47.6% (Jul 2026) | ACTIVE |
| LIBERO | 130 sim tasks | UT Austin (orig.) | \~97–98% (π0.5, MolmoAct 2) — saturated | DORMANT/saturated |
| ManiSkill3 | GPU sim tasks | UCSD/Hillbot | — | ACTIVE |
| HumanoidBench | 27 humanoid sim tasks | UC Berkeley | — | likely DORMANT |
| World Humanoid Robot Games | Athletics/skills | Beijing | 100 m in 9.39 s (Aug 2026, reported) | ACTIVE |
| RoboCup Humanoid | Soccer | RoboCup Federation | B-Human (Middle), Invic (Small), Tsinghua Hephaestus (Large) | ACTIVE |

---

## Gaps and pain points: data-economy opportunities, with evidence

1. **Commercial-rights gap.**
   - The largest and best real-robot and egocentric sets are non-commercial or no-derivatives: AgiBot World (CC BY-NC-SA), Galaxea, ActionNet, EgoDex (BY-NC-ND), LAFAN1 (BY-NC-ND), Motion-X and AMASS (research).
   - Commercially usable sets are older (OXE, Bridge), simulated (NVIDIA, Axis), gated (RoboMIND, Build AI) or video-only.
   - Luel (a rights-cleared marketplace) shows the angle.
   - → **Opportunity:** rights-cleared data with provenance and consent records, sold under commercial licenses.
2. **Quality and "training-ready" gap.**
   - A trajlens audit of 100 public LeRobot datasets (2026-06-29, tool-author blog): **81% had issues or failed to load cleanly** (47% errored, 21% timed out). **18.8% of those that linted successfully showed v2.1→v3.0 conversion corruption**, with episode boundaries mismatched to frames.
   - The Sep 2025 crawl: median 10 episodes, 43% with 1–5 episodes, 91% single-task.
   - → **Opportunity:** validation and QA, dedup, calibration checks, and selling "accepted/training-ready hours" (DataX Power reports 1.33–1.43× cost at 70–75% acceptance).
3. **Egocentric volume without actions.**
   - Egocentric-100K has 100k h but only 456×256 video, with no hand pose, depth or IMU.
   - Flikforge (a vendor) says raw footage resells at $2–5/h and buyers discard about 90%.
   - EgoVerse finds human data helps only when it is aligned with the robot objective.
   - Paired hand pose exists (EgoDex, EgoSuite, Open-AoE) but is either non-commercial or estimated rather than measured.
   - → **Opportunity:** calibrated, multi-sensor egocentric capture (stereo/depth, IMU, measured hand pose from gloves, wrist cameras), with coverage designed for a target skill rather than raw volume.
4. **Tactile and force data is scarce.**
   - Only RH20T (force/torque, plus tactile on 1 of 7 configurations), the AgiBot World visuo-tactile subset, Humanoid Everyday (tactile), and the new TacVerse visuo-tactile set.
   - LeRobot v0.6 added depth natively; tactile is supported only through third-party sensor plugins (README).
   - Commercial activity: MANUS haptic gloves are used by BrainCo, Shadow and others; Genesis AI built an e-skin glove but keeps it closed.
   - → **Opportunity:** tactile-glove and visuo-tactile datasets in a standardized LeRobot schema.
5. **Humanoid whole-body and loco-manipulation data is small or simulated.**
   - Unitree's open sets are about 40–120 episodes per task.
   - GRAIL is about 55 h and simulated.
   - ActionNet is about 140 h; Humanoid Everyday has 10.3k trajectories.
   - The BEHAVIOR robot is wheeled and in simulation.
   - Bones Studio shows mocap feeds humanoid controllers (BONES-SEED → NVIDIA SONIC).
   - → **Opportunity:** real whole-body teleoperation and retargeted mocap with contact labels.
6. **Long-horizon tasks remain unsolved.**
   - Best BEHAVIOR q-score is 26% (binary success about 12%).
   - RoboChallenge multi-step tasks (for example, sandwich making) are near 0%.
   - RoboCerebra and RMBench appeared in 2026.
   - → **Opportunity:** long-horizon demonstrations with subtask and stage annotations. LeRobot v0.6 added language columns and a VLM annotation pipeline, so the schema exists.
7. **Failure, recovery and intervention data.**
   - Few open sources: the RoboArena dumps (MIT; include scored rollouts) and XDOF's ABC (100 h of evaluations).
   - LeRobot v0.6 added "record eval rollouts as datasets", DAgger handover and HIL-SERL interventions.
   - → **Opportunity:** a failure-labeled corpus and recovery demonstrations.
8. **Evaluation bottleneck.**
   - LIBERO is saturated (96–98%).
   - Real-robot evaluation is costly. RoboChallenge runs a 10-robot fleet; RoboArena gathered 600+ pairwise episodes at launch.
   - ArmnetBench shows $359 cells with about 10 s of operator time per rollout.
   - Third-party evaluators (Cortex AI) now appear in model launches, but organizer conflicts of interest exist (Dexmal's DM0 leads RoboChallenge).
   - → **Opportunity:** neutral evaluation-as-a-service and evaluation logs sold as data.
9. **Price compression and labor arbitrage.**
   - Collector pay in India is $1–4.20/h (Human Archive and competitors), while list prices are $15–60 per raw hour.
   - DreamVu: "a real market, and it filled up fast."
   - Volume will commoditize; differentiation must come from sensors, QA, rights and robot-alignment.
10. **Format fragmentation.**
    - LeRobot v3 dominates the Hub (77k datasets), but raw logs are MCAP/rosbag, legacy data is RLDS/HDF5/WebDataset/VRS, and v2.1→v3 conversion is buggy.
    - any4lerobot (1.1k stars) shows the demand.
    - EgoSuite ships both LeRobot v3 and MCAP; LeRobot added Foxglove.
    - → **Opportunity:** robust MCAP↔LeRobot pipelines, schema validation, and Lance/streaming backends.
11. **Open data is orders of magnitude below closed data.**
    - Generalist reports 270k h, growing by 10k h per week, versus about 3k h for AgiBot World and 350 h for DROID.
    - Frontier labs buy from XDOF, Scale and Lightwheel.
    - → Open datasets work mainly as marketing and standards-setting for suppliers (Build AI, Lightwheel, Axis, XDOF all open-sourced data in 2026).

---

## Plain list of repos and datasets referenced, with counts fetched

**GitHub**

| Repo | Counts |
|---|---|
| huggingface/lerobot | 27k stars, 5.7k forks, 976 open issues; v0.6.1 (2026-08-03); PyPI 238k/month |
| ros2/ros2_documentation | — |
| ros2/rmw_zenoh | 497 stars, 111 forks |
| moveit/moveit2 | \~2.0k stars, 785 forks |
| ros-controls/ros2_control | 935 stars, 455 forks, 232 tags |
| NVIDIA-ISAAC-ROS/isaac_ros_common | 319 stars, 227 forks |
| dora-rs/dora | 3.9k stars, 436 forks; 1.0.1 (2026-09-03) |
| copper-project/copper-rs | 1.5k stars, 106 forks; v1.2.2 (2026-10-02) |
| eclipse-zenoh/zenoh | 3.0k stars, 350 forks; 1.10.1 |
| eclipse-zenoh/zenoh-python | 173 stars |
| rerun-io/rerun | 11k stars, 849 forks; 0.38.1 |
| foxglove/mcap | 1.1k stars, 236 forks |
| foxglove/foxglove-sdk | 311 stars, 110 forks |
| lichtblick-suite/lichtblick | 1.1k stars, 751 forks |
| viser-project/viser | 2.7k stars, 215 forks; 837k PyPI/month |
| robot-descriptions/robot_descriptions.py | 833 stars, 75 forks |
| phospho-app/phosphobot | 395 stars, 84 forks |
| robotology/yarp | 603 stars, 217 forks |
| Tavish9/any4lerobot | 1.1k stars, 103 forks |
| voxel51/fiftyone | 11k stars, 826 forks |
| lancedb/lancedb | 11k stars, 1.0k forks |
| berkeleyautomation/robodm | 165 stars, 21 forks |
| wuphilipp/gello_software | 538 stars, 178 forks |
| real-stanford/universal_manipulation_interface | 1.6k stars, 289 forks |
| haosulab/ManiSkill | 3.3k stars, 550 forks |
| carlosferrazza/humanoid-bench | 799 stars, 130 forks |
| Lifelong-Robot-Learning/LIBERO | 2.3k stars, 476 forks |

Also referenced, counts not fetched:

- google-deepmind/open_x_embodiment
- droid-dataset/droid
- OpenDriveLab/AgiBot-World
- apple/ml-egodex
- facebookresearch/Ego4d
- facebookresearch/hot3d
- SimarKareer/EgoMimic
- ubisoft/ubisoft-laforge-animation-dataset
- IDEA-Research/Motion-X
- lijiaman/omomo_release
- StanfordVL/BEHAVIOR-1K
- mli0603/openpi-comet
- robo-arena/roboarena
- RoboTwin-Platform/RoboTwin
- Spirit-AI-Team/spirit-v1.5 (per press)

**Hugging Face** (downloads are "last month")

| ID | Counts |
|---|---|
| tag `LeRobot` | 77,473 datasets |
| agibot-world/AgiBotWorld-Beta | 94,323 downloads; 80 likes |
| agibot-world/AgiBotWorld-Alpha | — |
| builddotai/Egocentric-100K | 173,756 downloads; 143 likes |
| builddotai/Egocentric-10K | 353 downloads |
| builddotai/Egocentric-100K-Evaluation | 255 |
| builddotai/Egocentric-10K-Evaluation | 293 |
| nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim | 1,267,768; 274 likes |
| nvidia/PhysicalAI-Robotics-Open-H-Embodiment | 121,469 |
| nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL | 40,234; 24 likes |
| x-humanoid-robomind/RoboMIND | 53,240; 54 likes |
| OpenGalaxea/Galaxea-Open-World-Dataset | 7,158; 53 likes |
| FourierIntelligence/ActionNet | — |
| unitreerobotics (org) | 162 datasets; unitreerobotics/G1_Dex1_HangCup 1,089 |
| TacVerse/opendata | \~6.1k (listing) |
| RoboArena/DataDump_02-03-2026 | 1,018 |
| lerobot/droid_1.0.1 | 24 likes |
| cadene/droid | 29 likes |
| physical-intelligence/libero | 43.3k (listing); 91 likes |
| behavior-1k/2026-challenge-demos | 3.27 TB |
| axisrobotics/Franka-Dataset | 160k+ (self-reported) |
| lerobot/robocerebra_unified | 6,660 episodes |
| EgoSteer/EgoSteer-RealWorld | 5.22k (listing) |

Also referenced, counts not fetched:

- LightwheelAI/EgoStandard (seen in search only)
- gatech/EgoMimic
- YuhongZhang/Motion-Xplusplus
- bop-benchmark/hot3d
- TianxingChen/RoboTwin2.0
- Spirit-AI-robotics/Spirit-v1.5
- nvidia/gr00t17-lerobot-libero_* (models)

---

### Other source URLs opened today (not listed per entry above)

- https://techcrunch.com/2026/06/17/collecting-robot-training-data-is-dirty-unglamorous-work-some-ai-labs-are-already-paying-xdof-to-do-it/
- https://techcrunch.com/2026/05/26/human-archive-taps-into-indias-services-startups-to-collect-data-for-physical-ai/
- https://www.dreamvu.ai/blog/robot-training-data-companies-2026
- https://www.exylos.ai/blog/robotics-dataset-pricing/
- https://www.dataxpower.com/blog/humanoid-robot-data-collection-cost
- https://flikforge.com/robot-data-glut-egocentric-video-physical-ai/
- https://generalistai.com/blog/nov-04-2025-GEN-0
- https://www.manus-meta.com/blog/the-emerging-data-challenge-in-dexterous-robot-hands-from-icra-2026
- https://www.prnewswire.com/news-releases/genesis-ai-unveils-gene-26-5--the-first-ai-brain-to-enable-robots-with-human-level-physical-manipulation-capabilities-302763638.html
- https://dev.to/kunalsomani/we-linted-100-public-lerobot-datasets-heres-what-we-found-3no0

### Facts that are headline-only (not opened)

- Foxglove $40M Series B (Nov 2025)
- Rerun $17M seed (TechCrunch); the SEK 170M figure itself was verified via Techleap
- MoveIt Pro 9 dates
- Isaac ROS 4.1
- LeRobot "58,000 datasets" (TechTimes)
- June 2025 hackathon numbers (LeRobot X post)
- NVIDIA×HF LeRobot article (July 2026)
- Midcentury, vision lab and Proception rounds
