Files: MD

Contents

Working notes for Open-Source Humanoid Robotics Landscape (Oct 2026), compiled Oct 2, 2026. Not fact-checked line by line: where these notes and the report disagree, trust the report. See README.md.

Open-source robot-learning tooling, open datasets and benchmarks: the data-economy slice (as of 2026-10-02)

How this was compiled. Read this first.

  • Sources and dates. Everything was fetched on 2026-10-02.
    • GitHub, PyPI and crates numbers come mostly from the GitHits package/repo index (refreshed 2026-09-15 to 2026-10-02), because direct GitHub and PyPI pages were rate-limited.
    • Hugging Face (HF) numbers come from HF pages opened today. HF "downloads last month" is a rolling file-request counter, not a count of unique users, and it is inflated for many-file or streamed datasets. Use it only to compare datasets with each other.
  • Tool failures (not retried):
    • Rate-limited (HTTP 429) early in the session:
      • github.com/huggingface/lerobot and /releases
      • pypi.org/project/lerobot
      • HF blog posts for LeRobot v0.5.0 and v0.6.0
      • HF docs page /docs/lerobot/lerobot-dataset-v3
      • techtimes.com article "LeRobot Hub surpasses 58,000 datasets"
      • letsdatascience LeRobot 0.6 article
      • roboticsandautomationnews NVIDIA×HF LeRobot article
      • arxiv.org/abs/2505.11709
      • egxodata.com "Robotics Data Release Tracker 2026"
      • behavior.stanford.edu 2025 leaderboard
    • Blocked: dexset.ai (robots.txt fetch timed out).
    • JavaScript-rendered pages with no data returned: robochallenge.ai, robo-arena.github.io, the RoboTwin leaderboard, the BEHAVIOR 2026 HF leaderboard Space, and xdof.ai.
    • Budgets exhausted: the shared WebSearch budget (200/session) and then the WebFetch budget (800/h) ran out near the end. RoboWorld (arXiv 2607.01060) and RobotArena∞ were not opened.
  • Labels used below:
    • "(headline only)" means the fact comes from a search-result title, not an opened page.
    • "(self-reported)" means a vendor or press claim that I did not verify independently.

Executive summary (for a founder supplying robot-learning data)

  1. LeRobot is now the de facto format and toolchain.
    • HF lists 77,473 datasets tagged LeRobot (today). A crawl in Sep 2025 found 16,065 public ones.
    • LeRobotDataset v3.0 shipped with lerobot ≥0.4.0 (Oct 2025). New large releases use it: BEHAVIOR-2026 demos (3.27 TB), EgoSuite-Open100K and RoboCerebra.
    • Most of the volume is tiny hobby data:
      • median of 10 episodes per dataset;
      • 43% of datasets have 1–5 episodes;
      • SO-100/SO-101/Koch/LeKiwi arms make up more than half;
      • 91% are single-task.
    • Quality is weak: one audit found 81% of 100 public datasets had issues.
  2. 2026 is the year of open egocentric human data.
    • Build AI Egocentric-100K: 100,405 h, Apache-2.0, but video only at 456×256.
    • Lightwheel EgoSuite-Open100K: 10k h live with 100k planned, hand pose, LeRobot v3 plus MCAP, and commercial training allowed.
    • EgoVerse: 1,362 h.
    • Ant Group Open-AoE: about 2,000 h from smartphones, MANO hand poses, CC BY 4.0.
    • Human Archive HA-Ego-500: 500 h, densely annotated.
    • Vendors themselves say raw footage is turning into a commodity: about $2–5/h resale, against $15–40/h list prices.
  3. The large real-robot open datasets are mostly non-commercial.
    • Non-commercial: AgiBot World (1.0M trajectories, 2,976 h), Galaxea (500 h), Fourier ActionNet (140 h), EgoDex (829 h), LAFAN1, Motion-X, AMASS.
    • Commercially usable sets are older (OXE, BridgeData V2), simulated (NVIDIA GR00T-sim 325k trajectories, Axis 50k), or gated (RoboMIND, Apache-2.0).
  4. Evaluation is moving from saturated simulation to real-robot, distributed and long-horizon tests.
    • LIBERO is saturated at 96–98%.
    • The BEHAVIOR 2025 winner scored a q-score of only 26%.
    • RoboChallenge's best is about 62% on Table30 V2 (per press).
    • Other real-robot efforts: RoboArena, the AgiBot World Challenge (526 teams) and ArmnetBench ($359 SO-101 cells).
    • Evaluation as a service is emerging as a niche (Cortex AI, RoboChallenge).
  5. Money is flowing into data suppliers.
    • Raises: XDOF $70M; Mecka about $60–68M; Genesis AI $105M seed (tactile data glove); Encord $60M Series C; Config $27M; Midcentury $15M (headline only); Axis $12M; Human Archive $8.2M.
    • Vendor-published prices: $15–60 per raw hour (egocentric or teleoperation) and $80–150/h for full humanoid multi-sensor capture.
    • Collector pay in India: $1–4.20/h.

Part A — Tooling and middleware

LeRobot (Hugging Face): end-to-end PyTorch robot-learning library, dataset format and Hub ecosystem

  • Status: ACTIVE.
    • Releases: v0.6.1 (2026-08-03), v0.6.0 (2026-07-06), v0.5.1 (2026-04-07), v0.5.0 (2026-03-09), v0.4.4 (2026-02-27), v0.4.3 (2026-01-22), v0.4.0 (2025-10-23).
    • Repo last pushed 2026-10-02.
  • Who:
    • Hugging Face robotics team (Paris).
    • Rémi Cadène (original lead) has left. He is CEO of UMA (Paris), building the "Northstar" humanoid. His co-founders are Simon Alibert (CTO, LeRobot co-founder), Robert Knight (Chief Robot Officer, SO-100 arm designer) and Pierre Sermanet (CSO, ex-DeepMind). Thomas Wolf is listed as an adviser to UMA (TNW, 2026-07-07). UMA's seed round size is unconfirmed (about $40M was reportedly sought).
    • Release notes are now dominated by @imstevenpmwork (who cuts the releases), @CarolinePascal, @pkooij, @Maximellerbach, @s1lent4gnt, @nicolas-rabault, @HaomingSong and @AdilZouitine.
    • ICLR 2026 paper authors include Cadène, Alibert, Capuano, Aractingi, Zouitine, Kooijmans, Pascal, Palma, Shukor, Aubakirova, Lhoest, Gallouédec and Wolf.
  • Links: GitHub huggingface/lerobot; docs huggingface.co/docs/lerobot; HF org lerobot; paper arXiv 2602.22818 (ICLR 2026).
  • Stats (today):
    • 27k stars, 5.7k forks, 976 open issues.
    • PyPI: 238k downloads/month, 12 versions published.
    • Contributor count not verified (GitHub page rate-limited).
  • License: Apache-2.0.
  • What it is, in detail:
    • Unified Robot/Teleoperator/Camera classes.
      • Native hardware: SO100/SO101, LeKiwi, Koch, HopeJR, OMX, EarthRover, Reachy2, gamepads, keyboards, phones, OpenARM, Unitree G1 (whole-body control added in v0.5.0) and the Seeed reBot B601.
      • Plugins are auto-discovered by package-name prefix (lerobot_robot_*, lerobot_teleoperator_*, lerobot_camera_*). Plugins exist for xArm, UR5e, Franka, AgileX Piper, WidowX, ARX5, I2RT YAM, GELLO, SpaceMouse, Quest, ROS 2 bridges, and tactile and depth cameras.
    • Policies in the README:
      • Imitation learning: ACT, Diffusion, VQ-BeT, Multitask DiT.
      • Reinforcement learning: HIL-SERL, TDMPC.
      • Vision-language-action models (VLAs): π0, π0-FAST, π0.5, GR00T N1.7 (replaced N1.5 in v0.6), SmolVLA, X-VLA, EO-1, MolmoAct2, WALL-OSS, EVO1.
      • World models: VLA-JEPA, LingBot-VA, FastWAM, LaWAM, FLUX 3 Action.
      • Reward models: SARM, TOPReward, Robometer.
    • LeRobotDataset v3.0 (lerobot ≥0.4.0):
      • Many episodes are packed per Parquet/MP4 file (v2 stored one file per episode).
      • Episode boundaries are resolved through Parquet metadata ("relational metadata").
      • StreamingLeRobotDataset streams directly from the Hub.
      • Optional LanceDB table backend.
    • Data-relevant additions in v0.6.x:
      • depth-map support, with depth units stored in the metadata;
      • separate RGB and depth codecs, plus libaom-AV1;
      • video re-encode and trim tools;
      • language columns replacing subtask_index;
      • a VLM subtask-annotation pipeline;
      • metadata-based episode filtering;
      • "record eval rollouts as LeRobot datasets";
      • DAgger smooth handover;
      • a 2× faster dataloader;
      • a lerobot-rollout command-line tool;
      • remote training on HF Jobs;
      • Foxglove visualization (Rerun was already supported);
      • an Isaac Teleop → SO-101 recording example.
    • Benchmarks wired in through EnvHub: LIBERO, MetaWorld, RoboCasa365, RoboTwin 2.0, RoboCerebra, RoboMME, LIBERO-plus and VLABench, with Docker smoke tests.
    • v0.6 breaking changes: the minimal pip install no longer includes dataset or training extras; the minimum PyTorch is 2.7; the RL stack was rebuilt (sac is now gaussian_actor).
    • v0.5.0: requires Python ≥3.12 and transformers v5.
  • Adoption / usage data:
    • 77,473 HF datasets tagged LeRobot (HF search page, today). Sep 2025 crawl: 16,065 public datasets (Kamenski). A TechTimes headline on 2026-05-25 said the Hub "surpasses 58,000 datasets in one year" (headline only).
    • June 2025 worldwide hackathon: 3,000+ participants, 44 countries, 250+ submissions (LeRobot X post, headline only). A Munich node provided 50+ SO-101 arms. I could not verify a 2026 worldwide edition.
    • NVIDIA partnership: a 2026-07-18 article headline, "Nvidia and Hugging Face expand LeRobot…" (headline only). The release notes back this up with GR00T N1.7 integration, nvidia/gr00t17-lerobot-libero_* checkpoints and the Isaac Teleop example.
    • HF co-runs RoboChallenge with Dexmal, and Lightwheel shipped EgoSuite in partnership with HF.
  • Benchmarks / results (LeRobot docs):
    • LIBERO: π0.5 97.5% average (97.0 / 99.0 / 98.0 / 96.0) versus OpenPI's own 96.85%.
    • GR00T N1.7: 96.5% average (preliminary, ≥50 episodes per suite).
    • LaWAM: 98.4 / 99.6 / 98.0 on spatial / object / goal.
  • Trade-offs vs alternatives:
    • It is learning-centric and Python-first. It is not a real-time middleware: deterministic, low-latency control still needs ROS 2, dora or Copper.
    • It changes fast, with breaking renames even in v0.6.1 (lerobot.types → lerobot.lerobot_types).
    • Converting from v2.1 to v3.0 has corrupted episode boundaries in the wild (18.8% of successfully linted datasets in one audit, see Gaps).
    • Alternatives: openpi (Physical Intelligence models and recipes), and RLDS/TFDS (legacy, dormant).
  • Sources:

ROS 2: the standard robot middleware and distribution

  • Status: ACTIVE. Lyrical Luth reached general availability on 2026-05-22. It is the 12th ROS 2 release, an LTS supported to May 2031, and its ROS Boss is Shane Loretz.

  • Who: Open Robotics/OSRF and the community. RMW vendors include eProsima (Fast DDS, the default), ZettaScale (Zenoh), RTI, Eclipse Cyclone DDS and GurumDDS.

  • Links: docs.ros.org; GitHub ros2/ros2_documentation.

  • Supported distros:

    DistroReleasedEnd of life
    Lyrical2026-05-22May 2031
    Kilted2025-05-23Dec 2026
    Jazzy2024-05-23May 2029
    Humble2022-05-23May 2027
  • License: core packages are Apache-2.0 (not re-verified today).

  • What is new:

    • Lyrical Tier-1 platforms are Ubuntu 26.04 "Resolute" (amd64/aarch64) and RHEL 10. Ubuntu Noble and Debian Trixie are Tier 3. The default RMW is still rmw_fastrtps_cpp.
    • Lyrical adds rosidl::Buffer zero-copy publishing: uint8[] fields become rosidl::Buffer<uint8_t> with pluggable backends (for example GPU/CUDA). It works with rmw_fastrtps_cpp and rmw_zenoh_cpp, and is the basis of Isaac ROS 5.0 dropping NITROS.
    • rmw_zenoh_cpp became Tier 1 and ships in the binaries from Kilted onward.
    • MCAP has been the default rosbag2 format since Iron (2023).
  • Adoption: not re-measured today.

  • Trade-offs: mature ecosystem with drivers, MoveIt and Nav2, but heavy for Python/ML workflows. Under rclpy, dora claims 10–17× lower latency (self-reported). The 5-year LTS cadence suits product companies.

  • Sources: via GitHits github.com/ros2/ros2_documentation (source/Releases.rst, Releases/Release-Lyrical-Luth.rst, lyrical/release-timeline.rst, lyrical/supported-platforms.rst, Release-Kilted-Kaiju.rst, About-Different-Middleware-Vendors.rst, Release-Iron-Irwini.rst).

MoveIt 2 / MoveIt Pro (PickNik): motion planning

  • Status: ACTIVE.
    • Qualcomm agreed to acquire PickNik (announced 2026-09-23; terms undisclosed). Qualcomm pledged to keep MoveIt 1 and 2 open source under their existing licenses, hardware-agnostic and with community roadmaps.
    • MoveIt Pro 9 shipped around April 2026, and 9.4.1 is dated 2026-07-07 (release-note URLs, headline only).
  • Who: PickNik Robotics. Dave Coleman (CPO) is quoted; on the Qualcomm side, Nakul Duggal (EVP).
  • Links: GitHub moveit/moveit2; docs.picknik.ai.
  • Stats: about 2.0k stars, 785 forks. Binary builds exist for Rolling, Lyrical, Jazzy and Humble, with stable branches for humble, jazzy and kilted.
  • License: MoveIt 2 is BSD-3-Clause. MoveIt Pro is proprietary.
  • What it is: an open-source planning, IK and collision stack. MoveIt Pro is a commercial runtime with behavior trees and perception-to-motion features (Google was the first customer).
  • Trade-offs: the de facto open planner, but new ownership by a chip vendor is a strategic risk to watch (this is my inference, not reported). The GPU alternative is cuMotion (Isaac ROS).
  • Sources: https://www.therobotreport.com/qualcomm-acquires-picknik-robotics-keep-moveit-open-source/ ; https://github.com/moveit/moveit2

ros2_control: hardware abstraction and controller framework for ROS 2

  • Status: ACTIVE. Branches exist for humble, jazzy and kilted; master serves Rolling and Lyrical.
  • Who: ros-controls community (PickNik, Stogl Robotics and others).
  • Links: GitHub ros-controls/ros2_control; control.ros.org.
  • Stats: 935 stars, 455 forks, 232 tags.
  • License: Apache-2.0.
  • What it is: a real-time controller manager plus hardware interfaces. It is the standard way to put a new arm or hand behind ROS 2.
  • Trade-offs: the ROS-native choice. LeRobot uses its own Python robot classes instead, with optional ROS 2 bridges.
  • Sources: https://github.com/ros-controls/ros2_control

NVIDIA Isaac ROS: GPU-accelerated ROS 2 packages (perception, cuMotion)

  • Status: ACTIVE.
    • 5.0.0 (2026-09-21): migrated to Lyrical and removed NITROS in favor of native rosidl::Buffer with a CUDA backend.
    • 4.6.0 (2026-08-18): Jetson Orin support, Isaac Sim 6.0, and cloud-control workflows for the Unitree G1.
    • 4.5.0 (2026-07-06): sunset of the GXF implementation, a DNN stereo decoder, cuMotion 1.1.0.
    • 4.1 (2026-02-02, headline only).
  • Who: NVIDIA.
  • Links: GitHub NVIDIA-ISAAC-ROS/isaac_ros_common; nvidia-isaac-ros.github.io.
  • Stats: isaac_ros_common has 319 stars, 227 forks. Its last update was 2026-09-21 (GPU SM partitioning with CUDA MPS).
  • License: Apache-2.0 for the packages (not re-verified). It depends on NVIDIA's CUDA/TensorRT stack.
  • Trade-offs: the best Jetson Thor/Orin path, but tied to NVIDIA hardware and its frequent architecture churn (GXF, then NITROS, now Buffer).
  • Sources: https://nvidia-isaac-ros.github.io/releases/index.html ; https://github.com/NVIDIA-ISAAC-ROS/isaac_ros_common

Zenoh / rmw_zenoh: pub/sub/query protocol and ROS 2 middleware

  • Status: ACTIVE. Zenoh 1.10.1 (2026-09-07), 1.10.0 (2026-08-14), 1.9.0 (2026-04-10).
  • Who: Eclipse Foundation project, commercially backed by ZettaScale.
  • Links: GitHub eclipse-zenoh/zenoh, eclipse-zenoh/zenoh-python, ros2/rmw_zenoh.
  • Stats:
    • zenoh: 3.0k stars, 350 forks, 2.9M total crate downloads.
    • zenoh-python: 173 stars.
    • rmw_zenoh: 497 stars, 111 forks.
  • License: EPL-2.0 OR Apache-2.0. rmw_zenoh is Apache-2.0.
  • What it is: a brokerless or routed pub/sub layer that works well over WAN and Wi-Fi. dora uses it for shared memory and cross-machine links.
  • Trade-offs: better than DDS on lossy, multi-site networks and for discovery storms. Fast DDS remains the default and the more battle-tested option.
  • Sources: pkg_info crates:zenoh and pypi:eclipse-zenoh ; https://github.com/ros2/rmw_zenoh ; ROS docs as above.

dora-rs: low-latency Rust dataflow framework for robotics and AI

  • Status: ACTIVE. 1.0.0 shipped on 2026-09-02 and 1.0.1 on 2026-09-03.
    • The wire format and node APIs are frozen for all of 1.x.
    • 1.0 does not interoperate with 0.x.
  • Who: dora-rs community (originated by Xavier Tao and Philipp Oppermann; from memory).
  • Links: GitHub dora-rs/dora; blog docs/blog/2026-09-02-dora-1.0.md.
  • Stats: 3.9k stars, 436 forks; dora-cli has 58k total crate downloads.
  • License: Apache-2.0 (crate). The PyPI node-API package lists MIT.
  • What it is: a YAML-defined graph of nodes written in Rust, Python, C or C++.
    • Every message is an Apache Arrow array.
    • Messages ≥4 KB go over Zenoh shared memory, zero-copy.
    • There is a coordinator/daemon split, plus record/replay.
  • Benchmarks: "node-to-node latency 10–17× lower than rclpy" on identical Python workloads (self-reported; examples/ros2-comparison).
  • Trade-offs: much lighter than ROS 2 and ML-friendly through Arrow, but a smaller driver ecosystem.
  • Sources: via GitHits github.com/dora-rs/dora docs/blog/2026-09-02-dora-1.0.md ; pkg_info crates:dora-cli, pypi:dora-rs.

Copper-rs: deterministic Rust robotics runtime

  • Status: ACTIVE. 1.0.0 (2026-07-02); 1.2.2 (2026-10-02).
  • Who: Copper Robotics / copper-project.
  • Links: GitHub copper-project/copper-rs.
  • Stats: 1.5k stars, 106 forks, 49k crate downloads.
  • License: Apache-2.0.
  • What it is: a compile-time scheduled task graph with structured logging and replay (the project describes itself as an "OS for physical AI").
  • Trade-offs: strong determinism and replay for safety and debugging, but a young ecosystem.
  • Sources: pkg_info and pkg_changelog crates:cu29.

Rerun: multimodal, time-series logging, visualization and data platform

  • Status: ACTIVE. 0.38.1 (2026-09-16), with about 10 releases between July and September 2026 and 133 versions in total.
  • Who: Rerun (Stockholm). Seed of SEK 170M (about $17M) in March 2025, led by Point Nine. Customers include Meta, Google and Hugging Face (Dealroom/Techleap).
  • Links: GitHub rerun-io/rerun; rerun.io.
  • Stats: 11k stars, 849 forks.
  • License: MIT OR Apache-2.0.
  • What it is: an SDK (Python, Rust, C++) plus a viewer for 3D, video and time-series data, with growing database features. LeRobot's visualizer depends on rerun-sdk (bumped to <0.34 in v0.6).
  • Trade-offs: the best open viewer for learning datasets. Foxglove is stronger for ROS fleet observability.
  • Sources: pkg_info pypi:rerun-sdk ; https://finder.techleap.nl/news/feed/rerun-raises-170m-for-smart-robots

Foxglove / MCAP / Lichtblick: robotics observability, the log format, and the open fork

  • Status: ACTIVE.
    • foxglove-sdk 0.28.0 (2026-09-29).
    • MCAP Python 1.5.0 (2026-09-24).
    • Lichtblick 1.29.1 (2026-09-08).
  • Who:
    • Foxglove raised a $40M Series B (Nov 2025, headline only).
    • Lichtblick is led by BMW AG (its LICENSE says "Copyright 2024 BMW AG, 2021-2024 Foxglove").
  • Links: GitHub foxglove/mcap, foxglove/foxglove-sdk, lichtblick-suite/lichtblick.
  • Stats:
    • MCAP: 1.1k stars, 236 forks.
    • foxglove-sdk: 311 stars.
    • Lichtblick: 1.1k stars, 751 forks, 4.6k npm downloads/month.
  • License (status):
    • The Foxglove app is proprietary and has been closed since 2024.
    • MCAP and foxglove-sdk are MIT.
    • Lichtblick is the MPL-2.0 continuation of the open-source Foxglove Studio code ("open core").
  • What it is / adoption:
    • MCAP is a self-describing, indexed container. It has been the default rosbag2 format since ROS 2 Iron.
    • EgoSuite-Open100K ships MCAP alongside LeRobot v3, and LeRobot 0.6 added Foxglove visualization.
  • Trade-offs: MCAP is the raw-log lingua franca, while LeRobot v3 is the training format. A pipeline that converts MCAP to LeRobot is a natural product.
  • Sources: pkg_info pypi:mcap, pypi:foxglove-sdk, npm:@lichtblick/suite ; via GitHits lichtblick README and LICENSE.

Viser: web-based 3D visualization from Python

  • Status: ACTIVE. 1.1.1 (2026-09-15), 1.1.0 (2026-08-16).
  • Who: viser-project (Brent Yi et al.; originated in the nerfstudio ecosystem — from memory).
  • Stats: 2.7k stars; 837k PyPI downloads/month.
  • License: MIT.
  • What it is / trade-offs: an easy browser GUI for robot, scene and policy debugging. It is not a logging database (Rerun is).
  • Sources: pkg_info pypi:viser.

robot_descriptions.py: one-line import of about 100+ URDF/MJCF robot models

  • Status: ACTIVE.
    • 3.2.0 (2026-09-12) added SRDFs and 5 new descriptions.
    • 3.1.0 (2026-07-23) added "GENE.01".
    • 3.0.0 (2026-07-11) added SRDF support.
  • Who: Stéphane Caron and contributors (from memory).
  • Stats: 833 stars, 75 forks.
  • License: Apache-2.0.
  • Sources: pkg_info pypi:robot_descriptions.

phosphobot: no-code control, recording and VLA-training app for low-cost arms

  • Status: SLOWING. The last PyPI/GitHub release was 0.3.134 on 2025-10-22. The repo is not archived.
  • Who: phospho (Paris).
  • Stats: 395 stars, 84 forks; 4.7k PyPI downloads/month.
  • License: MIT.
  • What it is: supports SO-100/101, Koch, WX-250, Piper and Unitree Go2; teleoperation by keyboard, gamepad, leader arm or Quest; HF integration.
  • Trade-offs: superseded in practice by LeRobot plus LeLab.
  • Sources: https://github.com/phospho-app/phosphobot ; pkg_info pypi:phosphobot.

YARP: middleware for iCub/IIT humanoids

  • Status: DORMANT for releases (unverified). The releases page shows YARP 4.0.1 dated 2024-09-07 (the page summarizer has misread years before).
  • Who: IIT robotology.
  • Stats: 603 stars, 217 forks.
  • License: BSD-3-Clause, with optional LGPL/GPL components.
  • Trade-offs: niche, tied to the iCub/ergoCub ecosystem.
  • Sources: https://github.com/robotology/yarp ; https://github.com/robotology/yarp/releases

Open-source data management, curation, annotation and capture tools

  • any4lerobot (Tavish9/any4lerobot): 1.1k stars, MIT.
    • Converts OXE, AgiBot World, RoboMIND, LIBERO and RoboCasa to LeRobot, and LeRobot to RLDS.
    • Migrates LeRobot versions, including v2.1 ↔ v3.0.
    • Likely ACTIVE, because v3.0 support postdates Oct 2025.
  • LeRobot built-ins: delete, split, merge and add/remove features; recompute stats; re-encode and trim video; VLM subtask annotation; metadata filtering. LanceDB storage is documented as a backend.
  • LanceDB: 11k stars, 7.9M PyPI downloads/month, Apache-2.0.
  • FiftyOne (Voxel51): 11k stars, 108k downloads/month, Apache-2.0. Voxel51 publishes Egocentric-10K subsets in its dataset zoo.
  • RoboDM (berkeleyautomation/robodm): 165 stars, v0.1.0 about a year ago. DORMANT/SLOWING.
  • RLDS: 0.1.8, about 3 years old. DORMANT (legacy OXE stack).
  • trajlens: v0.4.0, Apache-2.0. A new, tiny LeRobot dataset linter with 16 checks.
  • Capture hardware/software:
    • GELLO (wuphilipp/gello_software): 538 stars, MIT. Supports I2RT YAM, Franka FR3/FER, UR and xArm, with gravity compensation. Its creator Philipp Wu is now CEO of XDOF.
    • UMI (real-stanford/universal_manipulation_interface): 1.6k stars, MIT. The core repo is largely static (likely DORMANT); the ecosystem lives on in forks.
  • Sources: https://github.com/Tavish9/any4lerobot ; https://github.com/wuphilipp/gello_software ; https://github.com/real-stanford/universal_manipulation_interface ; pkg_info pypi:lancedb, fiftyone, robodm, rlds, trajlens.

Part B — Open datasets

LeRobot community datasets on the HF Hub: the long tail of crowdsourced robot data

  • Status: ACTIVE. 77,473 datasets match other=LeRobot today.
  • Profile (Kamenski crawl of 16,065 public datasets, Sep 2025):
    • SO-100/101, Koch and LeKiwi make up more than half; about 296 robot-type strings appear.
    • 640×480 at 30 fps is used by about 72% of cameras.
    • Median 10 episodes per dataset (mean 273); about 43% have 1–5 episodes.
    • Median episode length is 17.4 s.
    • 91% are single-task.
    • The largest single dataset has 442,226 episodes.
  • Currently trending (listing, downloads):
    • TacVerse/opendata: 6.1k (visuo-tactile)
    • EgoSteer/EgoSteer-RealWorld: 5.22k
    • physical-intelligence/libero: 43.3k
    • LightwheelAI/leisaac-pick-orange
  • License: varies by uploader. Many datasets carry no license.
  • Trade-offs: enormous breadth but shallow, with inconsistent calibration and quality (see Gaps).
  • Sources: https://huggingface.co/datasets?other=LeRobot ; https://www.kamenski.me/articles/analyzing-lerobot-datasets-on-hugging-face

Open X-Embodiment (OXE): pooled cross-embodiment robot data

  • Status: DORMANT as a dataset (no new versions; RLDS/GCS). It is still widely used and has been converted to LeRobot.
  • Who: Google DeepMind and 21 institutions (34 labs).
  • Links: GitHub google-deepmind/open_x_embodiment; arXiv 2310.08864.
  • Size: 1M+ real trajectories, 22 embodiments, 60 datasets, 527 skills (160,266 tasks).
  • License: code Apache-2.0; "all other materials" CC-BY 4.0. Constituent datasets keep their own terms.
  • Benchmarks: RT-1-X +50% in the low-data regime; RT-2-X about 3× on emergent skills (paper).
  • Trade-offs: breadth, but heterogeneous, with low-resolution, short and mostly-gripper data.
  • Sources: https://robotics-transformer-x.github.io/ ; via GitHits OXE README.

DROID: large in-the-wild Franka teleoperation dataset

  • Status: DORMANT. Updates: language annotations in Dec 2024, improved calibrations for 36k episodes in Apr 2025. It remains the platform for RoboArena.
  • Who: a multi-lab consortium (13 institutions, 50 collectors).
  • Links: GitHub droid-dataset/droid; droid-dataset.github.io; HF port lerobot/droid_1.0.1.
  • Size: 76k trajectories, 350 h, 564 scenes, 86 tasks. 1.7 TB in RLDS; 8.7 TB raw stereo MP4.
  • Embodiment and modalities: Franka Panda; 2× ZED 2 plus a wrist ZED Mini (stereo), with 1,417 viewpoints calibrated; Quest 2 teleoperation; 3 language annotations for 95% of successful episodes.
  • License: the HF port is labeled Apache-2.0 (40M rows; 24 likes). The original release terms were not re-verified.
  • Trade-offs: the best "diverse real scenes" single-arm set; single embodiment; no tactile or force data.
  • Sources: https://droid-dataset.github.io/ ; https://huggingface.co/datasets/lerobot/droid_1.0.1 ; via GitHits docs/the-droid-dataset.md.

BridgeData V2: WidowX tabletop data

  • Status: DORMANT (2023).
  • Who: UC Berkeley RAIL.
  • Size: 60,096 trajectories (50,365 teleoperated plus 9,731 scripted), 24 environments, 13 skills, 640×480 multi-view (including one RGB-D camera), with language.
  • License: CC BY 4.0 (commercial use OK).
  • Sources: https://rail-berkeley.github.io/bridgedata/

AgiBot World (Alpha/Beta): the largest open real-robot manipulation set

  • Status: the dataset's last major release was Beta (2025-03-01). The ecosystem is ACTIVE (ICRA 2026 challenge, Genie Sim 3.0).
  • Who: AgiBot and OpenDriveLab.
  • Links: GitHub OpenDriveLab/AgiBot-World; HF agibot-world/AgiBotWorld-Beta, agibot-world/AgiBotWorld-Alpha; arXiv 2503.06669; GO-1 / GO-1 Air models.
  • Size:
    • Beta: 1,003,672 trajectories, 2,976.4 h, about 43.8 TB (README) or 48.1 TB (HF card); 100 robots; 200+ task types; 87 atomic skills.
    • Alpha: 92,214 trajectories (8.5 TB).
  • Modalities: multi-view RGB, depth, proprioception, force/torque. A subset has visuo-tactile sensors and 6-DoF dexterous hands; there are mobile dual-arm robots.
  • Format: WebDataset, convertible to LeRobot.
  • License: CC BY-NC-SA 4.0 (non-commercial), gated.
  • Adoption: 94,323 downloads last month, 80 likes.
  • Trade-offs: unmatched scale, but non-commercial and single-vendor hardware.
  • Sources: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta ; via GitHits AgiBot-World README.

RoboMIND: multi-embodiment teleoperation set from Beijing's humanoid innovation center

  • Status: SLOWING/unclear. Current version is v1.2; "V2.0 announced on ModelScope" with no date verified.
  • Who: X-Humanoid (Beijing Humanoid Robot Innovation Center) and collaborators. RSS 2025.
  • Links: HF x-humanoid-robomind/RoboMIND; arXiv (Dec 2024).
  • Size: 107k trajectories, 479 tasks, 96 object classes, 12.3 TB.
    • Franka: 52,926
    • UR5e: 25,170
    • Tien Kung humanoid: 19,152
    • AgileX Cobot Magic: 10,629
  • Format: HDF5.
  • License: Apache-2.0, gated (accept conditions).
  • Adoption: 53,240 downloads last month, 54 likes.
  • Trade-offs: commercially licensed and multi-embodiment, but with no depth in parts.
  • Sources: https://huggingface.co/datasets/x-humanoid-robomind/RoboMIND

Galaxea Open-World Dataset: mobile bimanual R1-Lite in homes, kitchens, retail and offices

  • Status: SLOWING/DORMANT. Published 2025-08-30; no 2026 update verified.
  • Who: Galaxea AI (G0 model; arXiv 2509.00576).
  • Size: 500+ h, 227 tasks, 2.87 TB.
  • Modalities: 4 RGB streams (head, head-right, two wrists), joints, end-effector, IMU, chassis and torso actions; bilingual subtask language.
  • Format: LeRobot v2.1 (AV1, 15 fps).
  • License: CC BY-NC-SA 4.0, gated.
  • Adoption: 7,158 downloads last month, 53 likes.
  • Sources: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset

RH20T: contact-rich, multimodal real-robot dataset

  • Status: DORMANT (ICRA 2024).
  • Who: Shanghai Jiao Tong University.
  • Size: 110k+ sequences, 147 tasks, 7 robot configurations, about 40 TB (resized versions about 15 TB and 892 GB).
  • Modalities: RGB-D, 6-DoF force/torque at 100 Hz, audio, fingertip tactile (configuration 7), IR, and paired human demonstration videos.
  • License: split. Scenes 1–5 are CC BY-SA 4.0 (RH20T-C, commercial OK); scenes 6–10 are CC BY-NC 4.0 (RH20T-NC).
  • Trade-offs: one of the few sets with force/torque, audio and tactile, but older robots and heavy to download.
  • Sources: https://rh20t.github.io/

Fourier ActionNet: humanoid upper-body teleoperation with dexterous hands

  • Status: SLOWING (2025 release; no 2026 update found).
  • Who: Fourier Intelligence and SJTU.
  • Links: HF FourierIntelligence/ActionNet; action-net.org.
  • Size: 30k+ trajectories, about 140 h. Robots: GR1-T1, GR1-T2, GR2; 6-DoF and 12-DoF hands; OAK-D cameras; egocentric (head) video plus state and action.
  • License: CC BY-NC-SA 4.0 for data, Apache-2.0 for code.
  • Sources: https://action-net.org/

Humanoid Everyday: diverse humanoid manipulation dataset

  • Status: SLOWING (arXiv 2510.08807, Oct 2025; the cloud evaluation portal is "coming soon").
  • Who: USC (Yue Wang) and Toyota Research Institute.
  • Size: 10.3k trajectories, 3M+ frames, 260 tasks at 30 Hz.
  • Modalities: RGB, depth, LiDAR, tactile and language. Robot models were not listed on the page; I believe they are Unitree G1/H1, but this is unverified.
  • License: not stated on the page.
  • Sources: https://humanoideveryday.github.io/

Unitree datasets on HF: G1 teleoperation datasets, updated continuously

  • Status: ACTIVE. New G1 datasets were pushed "minutes ago" today.
  • Who: Unitree Robotics.
  • Links: HF org unitreerobotics (162 datasets). Models: UnifoLM-WMA-0, UnifoLM-ER-1/Flow (4B), UnifoLM-VLA.
  • What: per-task G1 datasets with Dex1 grippers and BrainCo or Inspire hands, including whole-body teleoperation ("WBT") tasks. Head-stereo plus wrist cameras; LeRobot-style chunked MP4 and Parquet.
  • Example: G1_Dex1_HangCup has 44 episodes, 8.8 GB and 1,089 downloads/month.
  • License: not shown on the cards I opened.
  • Trade-offs: small per-task sets (about 40–120 episodes), but the most active humanoid-specific open source.
  • Sources: https://huggingface.co/unitreerobotics ; https://huggingface.co/datasets/unitreerobotics/G1_Dex1_HangCup

NVIDIA Physical AI datasets (GR00T and others): sim, teleoperation and synthetic humanoid data

  • Status: ACTIVE. 33 datasets match "PhysicalAI-Robotics".
  • Key items:
    • nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim: about 325k sim trajectories (240k GR1 humanoid tabletop, 72k arm kitchen, 9k bimanual Panda, plus others), 1.91 TB, CC-BY-4.0, 1,267,768 downloads last month, 274 likes, 27 models trained on it.
    • nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL (2026-06-03, arXiv 2606.05160):
      • 22,312 physics-validated humanoid-object-interaction motions for the Unitree G1 (29 body DoF and 14 hand DoF), about 55 h, 224 GB.
      • Generated from synthetic video → SMPL-X/4D reconstruction → retargeting → RL validation in Isaac Lab.
      • Apache-2.0 (bundled assets keep upstream licenses); 40,234 downloads last month.
    • nvidia/PhysicalAI-Robotics-Open-H-Embodiment (created Feb 2026):
      • surgical and ultrasound robotics, built with 30+ organizations;
      • 750 h, 120k trajectories, 4.54 TB, LeRobot v2.1, CC-BY-4.0;
      • 121,469 downloads last month.
  • Trade-offs: commercially usable and very large, but mostly simulated or synthetic, so the sim-to-real gap applies.
  • Sources: https://huggingface.co/datasets?search=PhysicalAI-Robotics&sort=downloads ; the three dataset pages above.

BEHAVIOR dataset (BEHAVIOR-1K challenge demos): long-horizon household data in simulation

  • Status: ACTIVE.
  • Who: Stanford Vision and Learning Lab. Teleoperation via JoyLo whole-body interface; 2026 data collected by Simovation.
  • Links: GitHub StanfordVL/BEHAVIOR-1K; HF behavior-1k/2026-challenge-demos.
  • Size:
    • 2025: 10,000 demos, 1,200+ h, 50 tasks.
    • 2026: 20,000 demos, 1,950 h, 100 tasks, 3.27 TB, LeRobot v3 (customized).
  • Modalities: RGB-D, segmentation, object states, proprioception and actions, skill/subtask annotations. The robot is the Galaxea R1 (wheeled, bimanual) in OmniGibson.
  • License: code MIT. The data/asset license was not verified.
  • Sources: via GitHits BEHAVIOR-1K docs/challenge/index.md, archive/2025/index.md, dataset.md.

EgoDex (Apple): dexterous egocentric video with 3D hand and upper-body poses

  • Status: SLOWING/DORMANT (May 2025 paper).
  • Who: Apple.
  • Links: GitHub apple/ml-egodex; arXiv 2505.11709.
  • Size: 829 h of 30 Hz 1080p video from Apple Vision Pro (ARKit), 194 tabletop tasks. Paired 3D head, upper-body and hand poses plus language. Splits: 725 h train, 7 h test, 97 h extra. Format: MP4 plus HDF5.
  • License: CC BY-NC-ND (non-commercial, no derivatives).
  • Adoption: used by H-RDT and OpenEgo; there is an HF viewer Space.
  • Trade-offs: the best paired hand-pose quality at scale, but the license blocks commercial training.
  • Sources: via GitHits apple/ml-egodex README.

Ego4D / Ego-Exo4D (Meta and consortium): general egocentric video

  • Status: SLOWING. Ego-Exo4D V2 and Ego4D v2.1 (Goal-Step) are the latest.
  • Size:
    • Ego4D: about 3,670 h (paper; not re-verified today).
    • Ego-Exo4D V2: 1,286.30 video hours (221.26 ego-hours), 5,035 takes, Aria plus GoPro, with 3D.
  • License: custom data license agreement (gated). The tooling code is MIT.
  • Trade-offs: large and diverse, but not manipulation-focused, with few action labels.
  • Sources: via GitHits github.com/facebookresearch/Ego4d README.

HOT3D (Meta): egocentric 3D hand and object tracking

  • Status: DORMANT (2024). It is still used in BOP and hand-tracking challenges.
  • Size: 833 min, 3.7M+ images, 19 subjects, 33 objects, from Aria and Quest 3. Optical mocap ground truth; MANO and UmeTrack hands; 6-DoF object poses.
  • License: the full set is under the HOT3D license agreement. HOT3D-Clips are on HF (bop-benchmark/hot3d); OpenEgo lists them as CC BY-SA 4.0 (I did not check the HF card).
  • Sources: https://facebookresearch.github.io/hot3d/ ; via GitHits hot3d README.

EgoMimic, and its successor EgoVerse: human egocentric data co-trained with robot data

  • EgoMimic:
    • Status: DORMANT (superseded).
    • GitHub SimarKareer/EgoMimic; small HF set gatech/EgoMimic (groceries, laundry, bowl tasks; Aria human plus robot HDF5).
  • EgoVerse (arXiv 2604.07607; 2026-04-08, revised 2026-07-07):
    • ACTIVE.
    • 1,362 h, 80,000 episodes, 1,965 tasks, 240 scenes, 2,087 demonstrators worldwide; standardized formats.
    • Lead author Ryan Punamiya plus 38 co-authors, academic and industry (I believe this is the Georgia Tech EgoMimic lineage, but that is unverified); egoverse.ai.
    • Finding: policies improve with more human data only when it is aligned with the robot objective.
    • Lightwheel's EgoSuite follows the EgoVerse consortium standards.
  • Sources: https://arxiv.org/abs/2604.07607 ; via GitHits EgoMimic README.

Build AI Egocentric-10K / Egocentric-100K: factory-worker egocentric video at massive scale

  • Status: SLOWING for the public set (last updated 2026-02-16).
  • Who: Build AI, which has raised about $15M (DreamVu; self-reported).
  • Links: HF builddotai/Egocentric-100K, builddotai/Egocentric-10K, plus evaluation splits.
  • Size (100K): 100,405 h; 2,010,759 clips; 10.8B frames; 456×256 at 30 fps; monocular head-mounted fisheye. Video only: no audio, depth, IMU or hand pose.
  • License: Apache-2.0, gated (contact information required).
  • Adoption: 173,756 downloads last month, 143 likes. The 10K set had 353 downloads.
  • Note: DreamVu says Build AI reached "1M hours (Apr 2026, 14,228 workers)". The worker count matches the 100K set itself (100,405 h ÷ 7.06 h per worker ≈ 14.2k), so treat the 1M-hour claim as unverified.
  • Trade-offs: volume leader with a commercial license, but low resolution and no action or pose labels.
  • Sources: https://huggingface.co/builddotai ; https://huggingface.co/datasets/builddotai/Egocentric-100K ; https://www.dreamvu.ai/blog/robot-training-data-companies-2026

Lightwheel EgoSuite-Open100K: egocentric human data with hand and body pose

  • Status: ACTIVE. Released 2026-08-26.
  • Who: Lightwheel (Chinese name 光轮 / Guanglun Intelligence), in partnership with HF.
  • Size: 10,000 h live; 100,000 h targeted. 15,000+ tasks and 15,000+ scenes across 7 categories and 128 scene types. Head-mounted cameras, plus a wrist camera in the "EgoPro" variant. Hand pose, full-body pose in some subsets, event-level semantics.
  • Format: LeRobot v3 (streamable) plus MCAP.
  • License: "academic research and commercial training" per the blog; per-card terms apply. One card surfaced in search as "Request access to EgoSuite-Open100K" (LightwheelAI/EgoStandard, not opened), which suggests gating.
  • Sources: https://huggingface.co/blog/LightwheelAI/egosuite-open100k

Open-AoE (Ant Group): smartphone-captured egocentric manipulation data and toolchain

  • Status: ACTIVE (July 2026; arXiv 2607.14183).
  • Size: about 2,000 h, 8,000+ atomic-action tasks, 500+ contributors, 400+ phone models.
  • Modalities: RGB, MANO 21-joint bimanual hand pose, 6-DoF camera trajectories, English action annotations, scene and object text.
  • License: CC BY 4.0. Hosted on GitHub, HF and ModelScope.
  • Trade-offs: cheap, scalable consumer capture with a commercial license, but no depth or tactile data and pose that is estimated, not measured.
  • Sources: https://arxiv.org/html/2607.14183

HA-Ego-500 (Human Archive): densely annotated workplace egocentric data

  • Status: ACTIVE (Aug 2026).
  • Size: 500 h (519.3 h collected), 310k+ labeled steps, 60+ work environments; median step about 4 s; hands engaged in 92% of steps. Custom multi-camera rigs.
  • License: not disclosed.
  • Sources: https://ego500.humanarchive.ai/

OpenEgo: a unified egocentric-manipulation corpus

  • Status: SLOWING (Sep 2025).
  • Who: UT Dallas and Physical Automation.
  • Size: 1,107 h, 290 tasks, 600+ environments, merging 6 sources (CaptainCook4D, HOI4D, HoloAssist, EgoDex, HOT3D, HO-Cap) with standardized MANO-21 hands.
  • License: inherits each source's license (mixed, including NC and ND).
  • Sources: https://arxiv.org/html/2509.05513

AMASS: unified human mocap in SMPL

  • Status: DORMANT (ICCV 2019). Still foundational for humanoid retargeting.
  • Size: 40+ h, 300+ subjects, 11k+ motions, 15 mocap datasets.
  • License: registration plus the MPI license. It is a research license; commercial use requires separate terms (from memory; verify).
  • Sources: https://amass.is.tue.mpg.de/

LAFAN1 (Ubisoft): game-quality mocap

  • Status: DORMANT. Shot in 2017.
  • Size: 5 subjects, 77 sequences, 496,672 frames at 30 fps (about 4.6 h), BVH.
  • License: CC BY-NC-ND 4.0.
  • Note: widely retargeted to G1/H1 for locomotion RL. Retargeted derivatives inherit the non-commercial, no-derivatives terms (my reading of the license).
  • Sources: via GitHits ubisoft-laforge-animation-dataset README and license.txt.

OMOMO: full-body human-object interaction mocap

  • Status: DORMANT (2023).
  • Size: about 10 h (paper; not re-verified), 15 large objects (10 train, 5 test).
  • License: code MIT; it requires SMPL-H/SMPL-X, which are under non-commercial model licenses.
  • Sources: via GitHits lijiaman/omomo_release.

Motion-X / Motion-X++: whole-body (SMPL-X) motion with text

  • Status: DORMANT. Motion-X++ was reorganized on HF in Mar 2025.
  • License: non-commercial research only; access is by form. Original RGB videos are not redistributed.
  • Sources: via GitHits IDEA-Research/Motion-X README and LICENSE.

New 2026 robot-action datasets

  • Axis Sim Dataset V1 (Axis Robotics, 2026-09-04):
  • ABC / "ABC-130K" (XDOF with UC Berkeley BAIR, June 2026):
    • 130k manipulation trajectories, 300 h simulation, 100 h evaluations (TechCrunch).
    • DreamVu calls it the "largest open-source bimanual manipulation dataset".
    • HF ID and license not verified.
  • MolmoAct 2 data (AI2, 2026-05-05):
    • 720+ h bimanual YAM demonstrations, plus all training datasets and evaluation rollouts released; language labels grew from about 71k to 146k.
    • Results: LIBERO 97.2% (98.1% with depth reasoning); Franka real-world 87.1%; third-party Cortex AI evaluation score 0.51, first on 7 of 8 tasks.
    • License not verified.
    • Source: https://allenai.org/blog/molmoact2
  • TacVerse/opendata (Xense Robotics):
  • RoboArena data dump (2026-02-03):
  • RoboCerebra unified: lerobot/robocerebra_unified, LeRobot v3; 6,660 episodes, 571,116 frames, 1,728 language subtasks (long-horizon).

Part C — Benchmarks and competitions

BEHAVIOR Challenge (Stanford): long-horizon household tasks in OmniGibson

  • Status: ACTIVE. The 2026 edition launched 2026-07-02. Submission deadline 2026-10-16; winners announced 2026-11-04.
  • What is measured:
    • 2026: 100 full household tasks in 7 scenes (4 new). One track with RGB, depth and proprioception. Ranked by average task success with BDDL partial credit ("q-score"). Baselines are π0.5 and GR00T N1.7. Prize pool $11k ($5k / $3k / $2k / $1k open source).
    • 2025: NeurIPS 2025, 50 tasks, 18 teams from 4 countries.
  • 2025 results:
    • 1st, "Robot Learning Collective" (independent researchers Ilia Larchenko, Gleb Zarin, Akash Karnatak): q-score 26% on public and private test sets; binary success 11.2% (public) and 12.4% (private). Built on π0.5 with correlated-noise flow matching and stage tracking. Compute: about $13k (8×H200 for about 15 days plus fine-tuning).
    • 2nd, Team Comet (openpi-comet): q-score 0.2514 (held-out). Post-challenge they report 0.345 on public validation (self-reported).
  • Sources:

RoboArena: distributed, double-blind, pairwise real-robot evaluation on DROID

  • Status: ACTIVE. Latest public data dump 2026-02-03; the leaderboard site is live but JavaScript-rendered, so I could not read it.
  • Who: 27 authors across 7 academic institutions (arXiv 2506.18123, June 2025).
  • What is measured: crowd-sourced A/B preferences and progress scores of generalist policies on evaluator-chosen tasks, served remotely through a BasePolicy server.
  • Results: the paper used 600+ pairwise episodes across 7 policies. The current leaders were not verified (site unreadable).
  • Trade-offs: scalable and hard to game, but tied to a single platform (Franka/DROID).
  • Sources: https://arxiv.org/abs/2506.18123 ; https://huggingface.co/datasets/RoboArena/DataDump_02-03-2026 ; via GitHits robo-arena/roboarena README.

RoboChallenge (Dexmal and Hugging Face): large-scale real-robot evaluation as a service

  • Status: ACTIVE. Table30 V2 launched 2026-03-24; there is a CVPR 2026 workshop paper.
  • What is measured:
    • Fleet of 10 robots of 4 types (UR5, Franka, Cobot Magic Aloha, ARX-5).
    • Table30 has 30 tasks. Submission is by remote API: the model runs on the participant's side.
    • Up to 1,000 fine-tuning episodes per task.
    • Metrics: success rate plus progress score (0–10 per task over 10 rollouts).
    • V2 adds zero-shot and out-of-domain tests and a DOS-W1 mobile platform.
  • Results (snapshots, not directly comparable across versions):
    • Initial (Oct 2025): π0.5 43.7% success / 62.2 score; π0 28.3%; CogACT 11.7%.
    • Spirit v1.5 (Spirit AI) ranked #1 in Jan 2026, and its weights and code were then open-sourced.
    • Aggregator snapshot (Feb 2026): DM0 (Dexmal) 37.3% success / 49.08; π0.5 17.67%; π0 9%.
    • Table30 V2 (Mar 2026, Chinese press): DM0 62%, GigaBrain-0.1 about 52%, π0.5 42.67%, RDT-1B 15%. Multi-step tasks such as sandwich making are near 0%.
    • Conflict of interest: the organizer's own model (DM0) leads.
  • Sources:

AgiBot World Challenge: manipulation and world-model competition

RoboTwin 2.0: bimanual simulation benchmark and data generator

  • Status: ACTIVE.

    • ICML 2026.
    • RMBench (memory-dependent tasks) on 2026-03-03.
    • StarVLA support on 2026-02-20.
    • IsaacLab-Arena and RLinf support on 2026-01-23.
    • Integrated in LeRobot v0.6.
  • Who: RoboTwin-Platform team (Tianxing Chen et al.).

  • Links: GitHub RoboTwin-Platform/RoboTwin; arXiv 2506.18088; HF TianxingChen/RoboTwin2.0.

  • What is measured: 50 tasks on Aloha-AgileX. Train on 50 clean demos per task, evaluate 100 episodes per task. Easy (clean) vs Hard (domain-randomized) settings. 100k+ pre-collected trajectories, 731 objects, 5 embodiments.

  • Results (paper baselines):

    PolicyEasyHard
    π046.4%16.3%
    RDT34.5%13.7%
    DP355.2%5.0%
    ACT29.7%1.7%
    DP28.0%0.6%

    The live leaderboard was unreadable (JavaScript).

  • Sources: https://arxiv.org/html/2506.18088 ; via GitHits RoboTwin README ; https://robotwin-platform.github.io/leaderboard (protocol only).

ManiSkill3: GPU-parallel manipulation simulation and benchmark

  • Status: ACTIVE. mani-skill 3.0.1 (2026-04-21); repo pushed 2026-08-04.
  • Who: Hao Su lab (UCSD) / Hillbot. RSS 2025 paper.
  • Stats: 3.3k stars, 550 forks; 24k PyPI downloads/month.
  • License: Apache-2.0.
  • Results: 200k+ state-only FPS and 30k+ FPS with rendering on a single RTX 4090 (self-reported).
  • Sources: pkg_info pypi:mani_skill ; via GitHits ManiSkill docs.

LIBERO (brief): lifelong-learning simulation suites, now saturated

  • Status: the original repo is effectively DORMANT but remains the default sanity check.
  • Stats: 2.3k stars, 476 forks. 130 tasks in 4 suites. Code MIT, data CC BY 4.0.
  • Results: π0.5 96.85% (OpenPI) and 97.5% (LeRobot); GR00T N1.7 96.5%; MolmoAct 2 97.2–98.1%.
  • Successors: LIBERO-plus (robustness) and RoboCerebra.
  • Sources: https://github.com/Lifelong-Robot-Learning/LIBERO ; LeRobot docs via GitHits.

HumanoidBench: simulated humanoid locomotion and manipulation

  • Status: likely DORMANT. 45 commits total, no releases; the last commit date could not be verified.
  • Who: UC Berkeley (Sferrazza et al., RSS 2024).
  • Stats: 799 stars, 130 forks, MIT.
  • What is measured: 15 manipulation and 12 locomotion tasks (the page also cites 31 including variants) on a Unitree H1 with Shadow hands (G1 and Digit variants), using MuJoCo/MJX. Baselines: TD-MPC2, DreamerV3, SAC, PPO.
  • Sources: https://github.com/carlosferrazza/humanoid-bench

World Humanoid Robot Games: humanoid athletics and skills competition

  • Status: ACTIVE. The 2nd edition ran 2026-08-22 to 08-26 in Beijing.
  • Results (as reported by the Beijing municipal government; event rules not verified):
    • 666 teams, 16 countries, 2,000+ robots — +138% teams and 4× robots versus 2025.
    • 100 m in 9.39 s (Tianzhuo team; 21.5 s in 2025).
    • High jump 2.8843 m (Tiangong).
    • Long jump over 7 m.
    • Football moved from 3v3 to 5v5.
  • Sources: https://english.beijing.gov.cn/beijinginfo/sci/latesttrends/202608/t20260825_4836357.html

RoboCup Humanoid League: robot soccer

New 2026 benchmarks and evaluation efforts

  • ArmnetBench v0.1 (Armnet, arXiv 2607.24481, July 2026):
    • 4 SO-101 arms in 3 cells ($359 single-arm cell, $477 bimanual cell); 12 tasks; 7 policies; 2,518 rollouts.
    • Leader: π0.5 at 47.6% strict success (45.4% single-arm, 52.1% bimanual).
    • Source: https://arxiv.org/html/2607.24481v1
  • RoboCasa365: 365 kitchen tasks (about 65 atomic, about 300 composite).
  • RoboCerebra: long-horizon.
  • RoboMME, LIBERO-plus, VLABench: integrated into LeRobot v0.6.
  • RMBench: memory-dependent tasks, on top of RoboTwin.
  • Cortex AI: third-party real-robot evaluation, used by AI2 for MolmoAct 2.
  • Not opened (budget ran out): RoboWorld (neural-simulator policy evaluation, arXiv 2607.01060) and RobotArena∞ (real-to-sim benchmarking).

Part D — The data economy

Who supplies data (2026)

Unless marked otherwise, these entries are from DreamVu's 2026-08-06 landscape post. DreamVu is itself a vendor, so treat them as self-reported.

CompanyData / productFunding and scale (as reported)
XDOFTeleoperation pipelines on deployment robots, GELLO-style capture, planned egocentric wearables; ABC dataset$70M (Thrive, Spark, a16z, Lux, WndrCo); about 60 staff; 20 customers incl. "several frontier labs" (TechCrunch, 2026-06-17); founder Philipp Wu (GELLO)
Scale AITeleoperation and human demosCustomers: Physical Intelligence, Generalist, Cobot
LightwheelSimReady OpenUSD assets, EgoSuite, RoboFinalsCustomers named: Google DeepMind, Figure, AgiBot, ByteDance, Geely, BYD
Mecka AIBody-worn sensors and iPhones; full-body kinematicsAbout $68M total; claims a $100M run-rate in signed contracts; 1X is a customer (Flikforge says $60M in June 2026)
ConfigBimanual capture plus model$27M seed at a $200M valuation (Samsung ecosystem); targets 1M hours
Human ArchiveHead rigs with RGB-D, tactile gloves, mocap suits$8.2M (Wing, NVP, YC W26); 1,000+ active headsets; pays workers $1/h base, while competitors pay ₹250–400/h ($2.63–4.20) (TechCrunch, 2026-05-26)
Build AIFactory egocentric videoAbout $15M; open Egocentric-10K/100K
EncordCuration and annotation tooling (LiDAR, video)$110M total, including a $60M Series C (Feb 2026); physical-AI revenue up 10×
Bones StudioStudio optical mocapBONES-SEED, 142k+ sequences (used for NVIDIA SONIC)
LuelRights-cleared marketplace (incl. curated Ego4D/Ego-Exo4D) plus custom collectionLightspeed-backed
Axis RoboticsCrowdsourced browser teleoperation in simulation$12M seed (Hack VC)
Genesis AITactile e-skin data glove ("100× cheaper hardware, up to 5× collection efficiency" vs teleoperation; self-reported) plus egocentric and internet video plus sim$105M seed; nothing open-sourced (2026-05-06)
MANUSData gloves (Metagloves Pro Haptic)Users include ROBOTERA, BrainCo, TESOLLO, Xynova, Shadow Robot; retargeting for AgiLink, Inspire, Allegro (ICRA 2026 blog)
Also (headline only)Midcentury$15M seed
Also (headline only)vision lab (factory data)$6M seed
Also (headline only)Proception$11M seed
Also (headline only)Tacta Systems"large-scale skill capture"

Closed-data benchmark for scale: Generalist's GEN-0 was trained on 270,000 h of real-world manipulation data, growing by more than 10,000 h per week, collected through "data foundry partners" (Generalist blog, 2025-11-04). That is 2–3 orders of magnitude more than the largest open real-robot sets.

Price points (all vendor-published; units are not comparable — raw vs accepted vs training-ready hours)

Source (vendor, date)ItemPrice
DexSet (via EXYLOS, 2026-09-05)Egocentric video$15–40 per raw hour
DexSetTeleoperation$28–60 per raw hour
DexSetAnnotation+$8–25 per data hour
Robotics Center of Silicon ValleyProduction dataset with 500+ demos$50k–200k
Robotics Center of Silicon ValleyPer episode$8–35
DataX Power (2026-07-07)Single-arm teleoperation$15–30/h
DataX PowerBimanual ALOHA-style$40–80/h
DataX PowerEgocentric wearable$25–60/h
DataX PowerFull humanoid multi-sensor$80–150/h, or $50–150 per complex demo
DataX PowerQA overhead75% acceptance → 1.33× cost
DreamVuRobot-specific teleoperation$50–200/h; Build AI's open data is "effectively free"
Flikforge (CEO op-ed, 2026-07-07)Raw footage resale$2–5/h, versus $5–20/h to create; buyers discard about 90%

Formats that dominate:

  • LeRobotDataset (v2.1 moving to v3.0) for training and Hub distribution.
  • MCAP for raw ROS/robot logs.
  • Legacy and other formats: RLDS/TFDS (OXE), HDF5 (RoboMIND, EgoDex, ALOHA-style), WebDataset (AgiBot), VRS (Meta Aria), Zarr (UMI/Diffusion Policy).

Also notable (brief)

  • LeLab (huggingface/leLab), web UI for LeRobot — ACTIVE (added to the README in 2026).
  • Spirit v1.5 (Spirit AI), open VLA that ranked #1 on RoboChallenge in Jan 2026 — ACTIVE.
  • GR00T N1.7 checkpoints for LeRobot LIBERO (nvidia/gr00t17-lerobot-libero_*) — ACTIVE.
  • StarVLA, RLinf, XPolicyLab/RoboDojo: VLA/RL infrastructure that integrates RoboTwin and BEHAVIOR — ACTIVE (2026 updates).
  • Genie Sim 3.0 / EWMBench (AgiBot): simulation and world-model evaluation used in the ICRA 2026 challenge — ACTIVE.
  • Open-H-Embodiment: surgical/ultrasound LeRobot data from 30+ organizations — ACTIVE.
  • EgoSteer/EgoSteer-RealWorld: trending LeRobot-tagged dataset (5.22k downloads); contents not verified.
  • Reachy 2, HopeJR, SO-101, LeKiwi, OMX: HF-supported open hardware — ACTIVE.
  • UnifoLM-WMA/ER/VLA: Unitree's open models — ACTIVE.

Comparison tables

(a) Tooling

ProjectRoleLicenseStarsLatest releaseStatus
LeRobotLearning library, dataset format, HubApache-2.027kv0.6.1 (2026-08-03)ACTIVE
ROS 2Middleware/distroApache-2.0 (core)—Lyrical (2026-05-22)ACTIVE
MoveIt 2Motion planningBSD-3~2.0kbinaries for LyricalACTIVE (Qualcomm acquiring PickNik)
ros2_controlControllers/HALApache-2.0935232 tagsACTIVE
Isaac ROSGPU ROS packagesApache-2.0*319 (common)5.0.0 (2026-09-21)ACTIVE
Zenoh / rmw_zenohTransport/RMWEPL-2.0/Apache-2.03.0k / 4971.10.1 (2026-09-07)ACTIVE
dora-rsDataflow runtimeApache-2.03.9k1.0.1 (2026-09-03)ACTIVE
Copper-rsDeterministic Rust runtimeApache-2.01.5k1.2.2 (2026-10-02)ACTIVE
RerunLogging/visualizationMIT/Apache-2.011k0.38.1 (2026-09-16)ACTIVE
MCAPLog formatMIT1.1kpy 1.5.0 (2026-09-24)ACTIVE
Foxglove appObservabilityProprietary——ACTIVE (closed)
LichtblickOpen Foxglove forkMPL-2.01.1k1.29.1 (2026-09-08)ACTIVE
Viser3D web visualizationMIT2.7k1.1.1 (2026-09-15)ACTIVE
robot_descriptionsRobot modelsApache-2.08333.2.0 (2026-09-12)ACTIVE
phosphobotLow-cost arm applicationMIT3950.3.134 (2025-10-22)SLOWING
YARPiCub middlewareBSD-36034.0.1 (2024?, unverified)DORMANT?
any4lerobotFormat convertersMIT1.1k—likely ACTIVE
RLDSLegacy dataset libraryApache-2.0—0.1.8 (about 3 years ago)DORMANT

*Not re-verified.

(b) Datasets

DatasetSizeEmbodiment / captureModalitiesLicense (commercial use?)Year / last updateHF downloads per month
AgiBot World Beta1.0M trajectories / 2,976 hAgiBot fleet (100 robots), dexterous subsetRGB-D, force/torque, visuo-tactile subsetCC BY-NC-SA (no)202594,323
OXE1M+ trajectories22 embodimentsRGB, variedCC-BY 4.0 (yes; check sub-datasets)2023—
GR00T X-Emb Sim~325k trajectoriesGR1, Panda, G1 (sim)Sim RGB/stateCC-BY-4.0 (yes)20251,267,768
RH20T110k+ sequences7 arm configurationsRGB-D, force/torque, audio, tactilesplit: BY-SA / BY-NC2023–24—
RoboMIND107k trajectoriesFranka, UR5e, Tien Kung, AgileXRGB, stateApache-2.0 (yes; gated)v1.253,240
DROID76k trajectories / 350 hFrankaStereo RGB, languageHF port Apache-2.02024–25—
BridgeData V260,096 trajectoriesWidowXRGB-D, languageCC BY 4.0 (yes)2023—
Axis Sim V150k+ trajectoriesFR3 (sim)Sim"open" (unclear)2026-09160k+ (self-reported)
Fourier ActionNet30k+ trajectories / 140 hGR1/GR2 + 6/12-DoF handsEgocentric RGB, stateBY-NC-SA (no)2025—
BEHAVIOR 202620k demos / 1,950 hR1 (sim)RGB-D, segmentation, subtasksnot verified2026—
Humanoid Everyday10.3k trajectories / 260 tasksHumanoidRGB-D, LiDAR, tactilenot stated2025—
Galaxea500+ h / 227 tasksR1-Lite4×RGB, IMU, languageBY-NC-SA (no)20257,158
GRAIL22,312 motions / ~55 hG1 (sim)SMPL-X human-object interaction + G1Apache-2.0 (yes)2026-0640,234
Unitree HF162 datasetsG1 + Dex1/BrainCo/InspireStereo + wristnot shown2026 (daily)~1k per set
MolmoAct 2 YAM720+ hBimanual YAMRGB, languagenot verified2026-05—
TacVerse~1.28k episodesBimanualVisuo-tactileCC-BY-SA (yes, share-alike)2026~6.1k
Egocentric-100K100,405 hHuman head-mounted fisheyeVideo only, 456×256Apache-2.0 (yes; gated)2025 / Feb 2026173,756
EgoSuite-Open100K10k h (→100k)Human head (+wrist)Hand and body posecommercial training allowed2026-08—
Open-AoE~2,000 hSmartphonesMANO hands, camera poseCC BY 4.0 (yes)2026-07—
EgoVerse1,362 h / 80k episodesHuman egocentricAnnotationssee egoverse.ai2026—
Ego-Exo4D V21,286 hAria + GoProMulti-view, 3Dcustom agreement2024–25—
OpenEgo1,107 hMixed (6 sources)MANO-21mixed2025—
EgoDex829 hVision Pro3D hand/body poseBY-NC-ND (no)2025—
HA-Ego-500500 hMulti-camera rigDense step labelsnot disclosed2026-08—
HOT3D13.9 h (833 min)Aria / Quest 3Mocap ground truth for hands and objectsHOT3D agreement2024—
AMASS40+ hOptical mocapSMPLresearch license2019—
LAFAN14.6 hMocapBVHBY-NC-ND (no)2017/2020—

(c) Benchmarks

BenchmarkWhat is measuredOperatorCurrent leader and score (date)Status
BEHAVIOR50 → 100 long-horizon household tasks (sim), q-scoreStanford SVL2025: Robot Learning Collective, q 26%; 2026 results due 2026-11-04ACTIVE
RoboChallenge Table30 (V2)Real-robot fleet, success rate + progressDexmal + HFDM0 ~62% on V2 (Mar 2026, press); Spirit v1.5 #1 (Jan 2026)ACTIVE
RoboArenaPairwise real-robot preferences (DROID)7-institution consortiumNot verified (site JavaScript)ACTIVE
AgiBot World ChallengeR2A + world model; sim → G2 robot finalsAgiBotPrismBot (vivo) R2A; NeoVerse-ABot WM (Jun 2026)ACTIVE
RoboTwin 2.050 bimanual tasks, clean vs randomizedRoboTwin teamPaper: π0 46.4% / 16.3%ACTIVE
ArmnetBench12 tasks on SO-101 arm farmArmnetπ0.5 47.6% (Jul 2026)ACTIVE
LIBERO130 sim tasksUT Austin (orig.)~97–98% (π0.5, MolmoAct 2) — saturatedDORMANT/saturated
ManiSkill3GPU sim tasksUCSD/Hillbot—ACTIVE
HumanoidBench27 humanoid sim tasksUC Berkeley—likely DORMANT
World Humanoid Robot GamesAthletics/skillsBeijing100 m in 9.39 s (Aug 2026, reported)ACTIVE
RoboCup HumanoidSoccerRoboCup FederationB-Human (Middle), Invic (Small), Tsinghua Hephaestus (Large)ACTIVE

Gaps and pain points: data-economy opportunities, with evidence

  1. Commercial-rights gap.
    • The largest and best real-robot and egocentric sets are non-commercial or no-derivatives: AgiBot World (CC BY-NC-SA), Galaxea, ActionNet, EgoDex (BY-NC-ND), LAFAN1 (BY-NC-ND), Motion-X and AMASS (research).
    • Commercially usable sets are older (OXE, Bridge), simulated (NVIDIA, Axis), gated (RoboMIND, Build AI) or video-only.
    • Luel (a rights-cleared marketplace) shows the angle.
    • → Opportunity: rights-cleared data with provenance and consent records, sold under commercial licenses.
  2. Quality and "training-ready" gap.
    • A trajlens audit of 100 public LeRobot datasets (2026-06-29, tool-author blog): 81% had issues or failed to load cleanly (47% errored, 21% timed out). 18.8% of those that linted successfully showed v2.1→v3.0 conversion corruption, with episode boundaries mismatched to frames.
    • The Sep 2025 crawl: median 10 episodes, 43% with 1–5 episodes, 91% single-task.
    • → Opportunity: validation and QA, dedup, calibration checks, and selling "accepted/training-ready hours" (DataX Power reports 1.33–1.43× cost at 70–75% acceptance).
  3. Egocentric volume without actions.
    • Egocentric-100K has 100k h but only 456×256 video, with no hand pose, depth or IMU.
    • Flikforge (a vendor) says raw footage resells at $2–5/h and buyers discard about 90%.
    • EgoVerse finds human data helps only when it is aligned with the robot objective.
    • Paired hand pose exists (EgoDex, EgoSuite, Open-AoE) but is either non-commercial or estimated rather than measured.
    • → Opportunity: calibrated, multi-sensor egocentric capture (stereo/depth, IMU, measured hand pose from gloves, wrist cameras), with coverage designed for a target skill rather than raw volume.
  4. Tactile and force data is scarce.
    • Only RH20T (force/torque, plus tactile on 1 of 7 configurations), the AgiBot World visuo-tactile subset, Humanoid Everyday (tactile), and the new TacVerse visuo-tactile set.
    • LeRobot v0.6 added depth natively; tactile is supported only through third-party sensor plugins (README).
    • Commercial activity: MANUS haptic gloves are used by BrainCo, Shadow and others; Genesis AI built an e-skin glove but keeps it closed.
    • → Opportunity: tactile-glove and visuo-tactile datasets in a standardized LeRobot schema.
  5. Humanoid whole-body and loco-manipulation data is small or simulated.
    • Unitree's open sets are about 40–120 episodes per task.
    • GRAIL is about 55 h and simulated.
    • ActionNet is about 140 h; Humanoid Everyday has 10.3k trajectories.
    • The BEHAVIOR robot is wheeled and in simulation.
    • Bones Studio shows mocap feeds humanoid controllers (BONES-SEED → NVIDIA SONIC).
    • → Opportunity: real whole-body teleoperation and retargeted mocap with contact labels.
  6. Long-horizon tasks remain unsolved.
    • Best BEHAVIOR q-score is 26% (binary success about 12%).
    • RoboChallenge multi-step tasks (for example, sandwich making) are near 0%.
    • RoboCerebra and RMBench appeared in 2026.
    • → Opportunity: long-horizon demonstrations with subtask and stage annotations. LeRobot v0.6 added language columns and a VLM annotation pipeline, so the schema exists.
  7. Failure, recovery and intervention data.
    • Few open sources: the RoboArena dumps (MIT; include scored rollouts) and XDOF's ABC (100 h of evaluations).
    • LeRobot v0.6 added "record eval rollouts as datasets", DAgger handover and HIL-SERL interventions.
    • → Opportunity: a failure-labeled corpus and recovery demonstrations.
  8. Evaluation bottleneck.
    • LIBERO is saturated (96–98%).
    • Real-robot evaluation is costly. RoboChallenge runs a 10-robot fleet; RoboArena gathered 600+ pairwise episodes at launch.
    • ArmnetBench shows $359 cells with about 10 s of operator time per rollout.
    • Third-party evaluators (Cortex AI) now appear in model launches, but organizer conflicts of interest exist (Dexmal's DM0 leads RoboChallenge).
    • → Opportunity: neutral evaluation-as-a-service and evaluation logs sold as data.
  9. Price compression and labor arbitrage.
    • Collector pay in India is $1–4.20/h (Human Archive and competitors), while list prices are $15–60 per raw hour.
    • DreamVu: "a real market, and it filled up fast."
    • Volume will commoditize; differentiation must come from sensors, QA, rights and robot-alignment.
  10. Format fragmentation.
    • LeRobot v3 dominates the Hub (77k datasets), but raw logs are MCAP/rosbag, legacy data is RLDS/HDF5/WebDataset/VRS, and v2.1→v3 conversion is buggy.
    • any4lerobot (1.1k stars) shows the demand.
    • EgoSuite ships both LeRobot v3 and MCAP; LeRobot added Foxglove.
    • → Opportunity: robust MCAP↔LeRobot pipelines, schema validation, and Lance/streaming backends.
  11. Open data is orders of magnitude below closed data.
    • Generalist reports 270k h, growing by 10k h per week, versus about 3k h for AgiBot World and 350 h for DROID.
    • Frontier labs buy from XDOF, Scale and Lightwheel.
    • → Open datasets work mainly as marketing and standards-setting for suppliers (Build AI, Lightwheel, Axis, XDOF all open-sourced data in 2026).

Plain list of repos and datasets referenced, with counts fetched

GitHub

RepoCounts
huggingface/lerobot27k stars, 5.7k forks, 976 open issues; v0.6.1 (2026-08-03); PyPI 238k/month
ros2/ros2_documentation—
ros2/rmw_zenoh497 stars, 111 forks
moveit/moveit2~2.0k stars, 785 forks
ros-controls/ros2_control935 stars, 455 forks, 232 tags
NVIDIA-ISAAC-ROS/isaac_ros_common319 stars, 227 forks
dora-rs/dora3.9k stars, 436 forks; 1.0.1 (2026-09-03)
copper-project/copper-rs1.5k stars, 106 forks; v1.2.2 (2026-10-02)
eclipse-zenoh/zenoh3.0k stars, 350 forks; 1.10.1
eclipse-zenoh/zenoh-python173 stars
rerun-io/rerun11k stars, 849 forks; 0.38.1
foxglove/mcap1.1k stars, 236 forks
foxglove/foxglove-sdk311 stars, 110 forks
lichtblick-suite/lichtblick1.1k stars, 751 forks
viser-project/viser2.7k stars, 215 forks; 837k PyPI/month
robot-descriptions/robot_descriptions.py833 stars, 75 forks
phospho-app/phosphobot395 stars, 84 forks
robotology/yarp603 stars, 217 forks
Tavish9/any4lerobot1.1k stars, 103 forks
voxel51/fiftyone11k stars, 826 forks
lancedb/lancedb11k stars, 1.0k forks
berkeleyautomation/robodm165 stars, 21 forks
wuphilipp/gello_software538 stars, 178 forks
real-stanford/universal_manipulation_interface1.6k stars, 289 forks
haosulab/ManiSkill3.3k stars, 550 forks
carlosferrazza/humanoid-bench799 stars, 130 forks
Lifelong-Robot-Learning/LIBERO2.3k stars, 476 forks

Also referenced, counts not fetched:

  • google-deepmind/open_x_embodiment
  • droid-dataset/droid
  • OpenDriveLab/AgiBot-World
  • apple/ml-egodex
  • facebookresearch/Ego4d
  • facebookresearch/hot3d
  • SimarKareer/EgoMimic
  • ubisoft/ubisoft-laforge-animation-dataset
  • IDEA-Research/Motion-X
  • lijiaman/omomo_release
  • StanfordVL/BEHAVIOR-1K
  • mli0603/openpi-comet
  • robo-arena/roboarena
  • RoboTwin-Platform/RoboTwin
  • Spirit-AI-Team/spirit-v1.5 (per press)

Hugging Face (downloads are "last month")

IDCounts
tag LeRobot77,473 datasets
agibot-world/AgiBotWorld-Beta94,323 downloads; 80 likes
agibot-world/AgiBotWorld-Alpha—
builddotai/Egocentric-100K173,756 downloads; 143 likes
builddotai/Egocentric-10K353 downloads
builddotai/Egocentric-100K-Evaluation255
builddotai/Egocentric-10K-Evaluation293
nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim1,267,768; 274 likes
nvidia/PhysicalAI-Robotics-Open-H-Embodiment121,469
nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL40,234; 24 likes
x-humanoid-robomind/RoboMIND53,240; 54 likes
OpenGalaxea/Galaxea-Open-World-Dataset7,158; 53 likes
FourierIntelligence/ActionNet—
unitreerobotics (org)162 datasets; unitreerobotics/G1_Dex1_HangCup 1,089
TacVerse/opendata~6.1k (listing)
RoboArena/DataDump_02-03-20261,018
lerobot/droid_1.0.124 likes
cadene/droid29 likes
physical-intelligence/libero43.3k (listing); 91 likes
behavior-1k/2026-challenge-demos3.27 TB
axisrobotics/Franka-Dataset160k+ (self-reported)
lerobot/robocerebra_unified6,660 episodes
EgoSteer/EgoSteer-RealWorld5.22k (listing)

Also referenced, counts not fetched:

  • LightwheelAI/EgoStandard (seen in search only)
  • gatech/EgoMimic
  • YuhongZhang/Motion-Xplusplus
  • bop-benchmark/hot3d
  • TianxingChen/RoboTwin2.0
  • Spirit-AI-robotics/Spirit-v1.5
  • nvidia/gr00t17-lerobot-libero_* (models)

Other source URLs opened today (not listed per entry above)

Facts that are headline-only (not opened)

  • Foxglove $40M Series B (Nov 2025)
  • Rerun $17M seed (TechCrunch); the SEK 170M figure itself was verified via Techleap
  • MoveIt Pro 9 dates
  • Isaac ROS 4.1
  • LeRobot "58,000 datasets" (TechTimes)
  • June 2025 hackathon numbers (LeRobot X post)
  • NVIDIA×HF LeRobot article (July 2026)
  • Midcentury, vision lab and Proception rounds