Working notes for Open-Source Humanoid Robotics Landscape (Oct 2026), compiled Oct 2, 2026. Not fact-checked line by line: where these notes and the report disagree, trust the report. See README.md.
Open-weight robot foundation models: VLAs, generalist policies, world models and benchmarks (as of 2 Oct 2026)
Slice: open-weight vision-language-action models (VLAs), generalist policies, robotics world models and world-action models, and the benchmark results that compare them. Unless a line says otherwise, every number was fetched on 2026-10-02.
Access notes and caveats (read first)
- Search budget ran out. The session's shared WebSearch budget (200 calls, shared by all agents) was used up early. After that I worked only from direct page fetches of primary sources: GitHub repo pages, Hugging Face model cards and listings, arXiv, company blogs and PyPI.
- Fetch budget was rate-limited. The shared WebFetch budget (800/hour across agents) cut off several times.
- Pages I could not read, and why:
- Rate-limited (HTTP 429):
github.com/Physical-Intelligence/openpi(main README). I used the repo's /releases and /issues pages, the PI website and the LeRobot docs instead.huggingface.co/blog/nvidia/gr00t-n1-7huggingface.co/nvidia/GR00T-H-N1.7- The letsdatascience.com and buildfastwithai.com NVIDIA articles.
- The PR Newswire RoboChallenge release. I read Yahoo's syndicated copy instead.
arxiv.org/abs/2601.02456v2. The unversioned abs page loaded later.
- Fetch budget exhausted:
arxiv.org/abs/2605.02881. I used the HF papers page instead. - Blocked by robots.txt: the openpi /pulls page.
- HTTP 403:
pi.website/research/human_to_robot. - Wikipedia Skild AI page: cache-only.
- Rendered by JavaScript, so the fetcher saw an empty page: the live RoboArena, RoboChallenge and RoboTwin 2.0 leaderboards. Their numbers below come from the papers, model cards and one third-party aggregator snapshot.
- Rate-limited (HTTP 429):
- HF download counts are inconsistent. A model card's "Downloads last month" and the HF listing pages sometimes differ by 3–30×. Examples: openvla-7b shows 460k on the card and 169k on the listing; RDT2-VQ shows 371 on the card and 13.7k on the listing. I give the card figure first and the listing figure in parentheses. Treat all of them as order-of-magnitude.
- "(not re-verified)" means the figure comes from my prior knowledge of the primary paper and I could not re-open it today.
- "Self-reported" means the model's authors produced the number. "Independent" means a third party measured it.
Executive summary for a data / simulation / environments founder
-
Three hubs.
- Physical Intelligence's openpi: π0, π0-FAST, π0.5. 13.2k stars.
- NVIDIA Isaac GR00T: N1.7. 8.2k stars, and now under a commercial-use license.
- Hugging Face LeRobot: 27.8k stars, v0.6.1 (Aug 3, 2026). This is the distribution layer that wraps most other models.
- HF has 26,403 models tagged "robotics" and 23,950 tagged lerobot. The two tags overlap.
- SmolVLA has the widest fine-tuning footprint: 7,808 fine-tunes of
lerobot/smolvla_base, against 633 forlerobot/pi05_base.
-
The frontier labs stopped opening their best models.
- PI's 2025–26 models (π*0.6, the MEM memory work) are closed. openpi's last new open checkpoints were π0.5 in Sept 2025. In openpi issue #980 (Jun 19, 2026) users ask whether newer models will be open-sourced.
- NVIDIA's GR00T N2, a world-action model, was previewed Mar 16, 2026 and is "slated for… end of 2026".
- Google's Gemini Robotics 2 (Jul 30, 2026) is available only through the API, early-access partners or trusted testers.
-
Most 2026 open releases come from Chinese companies, plus AI2.
- Chinese companies: Xiaomi, Spirit AI, GigaAI, Dexmal, Ant/Robbyant, Galaxea, X Square, AgiBot, Unitree, Shanghai AI Lab.
- AI2: MolmoAct2.
- The common recipe is a 3–6B model: a Qwen-VL backbone plus a flow-matching DiT "action expert".
-
Licensing is a real bottleneck, and it is about data.
- NVIDIA staff (Jul 2025) said GR00T weights were research-only "due to pre-training data license constraints… it's a data issue, not a license choice." N1.7 fixed this in Apr 2026.
- Many open models are still non-commercial: AgiBot GO-1 and Genie Envisioner (CC BY-NC-SA), InternVLA-A1 (CC BY-NC-SA), DreamZero (CC-BY-NC), and Galaxea G0.5 onward.
-
The data race is moving toward human and egocentric data with shared action spaces. Pretraining corpora disclosed so far:
Model Pretraining data GR00T N1.7 20K h of "EgoScale" human video plus robot data, with a relative end-effector action space shared by humans and robots RDT2 10k+ h collected with UMI (Universal Manipulation Interface) handheld grippers LingBot-VLA 20k h of real-robot data GigaBrain-0.7 37k+ h OpenWAM ~6,400 h, 30% egocentric human MolmoAct2 720 h bimanual; AI2 calls it the "largest open bimanual dataset" TRI's LBM ~1,700 h -
Benchmarks are saturated or show brittleness.
- LIBERO averages are now 98–99% for at least seven models.
- LIBERO-Plus and LIBERO-PRO show collapse under changes in camera viewpoint or the robot's starting pose, and show that models ignore language.
- Independent real-robot evaluations are far lower:
- RoboChallenge Table30: the top model scores about 50–60% success; π0.5 trained as a multi-task generalist scores 17.7%.
- BEHAVIOR 2025 (long-horizon household tasks): the winner scored 26%.
-
World-action models are the 2026 trend. These use a video-generation backbone to predict future frames and actions together. Examples: DreamZero → GR00T N2, Cosmos 3 Nano-Policy, GE-Act v2, UnifoLM-WMA, OpenWAM, LingBot-VA and FastWAM. They need large amounts of action-labelled video and are heavy to run; DreamZero is 14B and runs at about 7 Hz.
A. Most-used generalist VLAs
Physical Intelligence openpi (π0, π0-FAST, π0.5) — the reference open generalist VLA family
- Status: SLOWING on the open-source side; still heavily used.
- The last new open checkpoints were π0.5 plus a PyTorch port in Sept 2025 (per PI and X announcements; README not re-verified).
- The repo has no GitHub releases or tags.
- Issues were still being opened through Jul 9, 2026; 239 are open.
- LeRobot ports were updated in Jun 2026.
- PI's later work is closed: π*0.6 (Nov 17, 2025, learns from experience via reinforcement learning) and "VLAs with Long and Short-Term Memory" (Mar 3, 2026).
- Who: Physical Intelligence, San Francisco.
- Founders: Karol Hausman, Sergey Levine, Chelsea Finn, Brian Ichter, Lachy Groom, Adnan Esmail, Quan Vuong.
- Funding (Wikipedia):
- 2024: $400M at $2.4B (Bezos, OpenAI, Thrive, Lux, Bond).
- 2025: $600M at $5.6B, led by CapitalG.
- 2026: about $1B at about $11B (Founders Fund, Lightspeed, Thrive, Index, NVIDIA, Bezos, T. Rowe Price).
- TechCrunch reported the $1B talks on Mar 27, 2026, and a Dealroom headline says "$1.6B across two rounds, valuation $11.2B". I saw these as search-result headlines only.
- Links:
- GitHub
Physical-Intelligence/openpi. - HF ports:
lerobot/pi0_base,lerobot/pi0fast-base,lerobot/pi05_base,lerobot/pi05_libero_base,lerobot/pi05_libero_finetuned_v044. - Papers: π0 arXiv 2410.24164; FAST 2501.09747; π0.5 2504.16054 (IDs from prior knowledge).
- GitHub
- Stats:
- GitHub: 13.2k stars, 2.3k forks.
lerobot/pi05_base: 19,354 downloads last month, 103 likes, 633 fine-tunes, 8 adapters.pi0_base: about 12–19k a month (listing).pi0fast-base: about 1.9k a month (listing).
- License:
- Code: Apache-2.0 (the LeRobot docs say this is "consistent with the original OpenPI repository").
- Weights are derived from PaliGemma. The LeRobot ports carry the "gemma" license tag, and the PaliGemma tokenizer is gated on HF. Gemma terms allow commercial use subject to Google's use policy; have counsel check.
- What it is:
- π0 (Oct 31, 2024):
- Architecture: a 3B PaliGemma VLM plus a flow-matching action expert, about 3.3B in total (paper, not re-verified). A 470M "π0-small" variant exists.
- Control rate: up to 50 Hz.
- Data: 7–8 embodiments (UR5e, bimanual UR5e, Franka, bimanual Trossen, bimanual ARX, mobile Trossen, mobile Fibocom) plus Open X-Embodiment (OXE). The paper cites about 10k hours (not re-verified).
- π0-FAST: an autoregressive variant that uses FAST, a compressed action-token scheme. It trains faster and follows language better, but its inference is slower.
- π0.5:
- Co-trained on web multimodal data, verbal instructions, subtask commands, cross-embodiment data, multi-environment data and about 400 h of mobile-manipulation data.
- The robot's state is discretized into 256 bins and written into the prompt.
- LeRobot executes 10-step action chunks by default.
- π0 (Oct 31, 2024):
- Compute:
- LeRobot recommends one 80 GB GPU to fine-tune π0.5 (with gradient checkpointing and bf16). Its LIBERO reproduction used 8×H100, batch 256, 6k steps.
- The openpi README lists >8 GB for inference, >22.5 GB for LoRA and >70 GB for full fine-tuning (not re-verified).
- Adoption:
- 633 HF fine-tunes of π0.5.
- Baseline in BEHAVIOR 2026 and RoboChallenge.
- The BEHAVIOR 2025 winning entry was built on π0.5.
- The openpi issues show deployments on SO-101 arms, ALOHA and DROID setups.
- Benchmarks:
- LIBERO, self-reported by openpi via the LeRobot docs: π0.5 scores 98.8 / 98.2 / 98.0 / 92.4 on Spatial / Object / Goal / Long, average 96.85.
- LIBERO, LeRobot re-run (10 episodes per task): 97.0 / 99.0 / 98.0 / 96.0, average 97.5.
- LIBERO for π0 and π0-FAST, as compiled in the Xiaomi-Robotics-0 paper:
- π0: 96.8 / 98.8 / 95.8 / 85.2, average 94.2.
- π0-FAST: 96.4 / 96.8 / 88.6 / 60.2, average 85.5.
- LIBERO-Plus (independent): π0 drops from 94.2 to 15.8 under camera changes, 6.6 under robot-start changes and 61.0 under language changes. π0-FAST is more camera-robust at 66.4.
- LIBERO-PRO (independent): π0 and π0.5 drop to about 0% when objects are moved.
- RoboChallenge Table30 (organizer-run, Oct 2025): π0.5 task-specific 43.7% success (score 62.2); π0 28.3% (47.6); multi-task generalists π0.5 17.7% and π0 9.3%.
- RoboArena paper (v2, Nov 29, 2025): π0-FAST-DROID ranked highest of 7 DROID policies.
- SimplerEnv (a simulated replica of Google Robot and WidowX setups) is not published by PI. Third parties cite π0 WidowX 58.8 (X-VLA table) and π0-FAST WidowX 48.3 (InternVLA-M1 table).
- Trade-offs:
- Strengths: the most evidence of real-world generalization and the largest community.
- Weaknesses:
- The codebase is JAX-first; issue #989 (Jun 30, 2026) reports "performance gap between PyTorch and JAX fine-tuning of pi0.5_base on LIBERO / LIBERO-Plus OOD".
- Users are asking how to export to TensorRT (#982).
- Data-conversion scripts are "outdated and extremely memory-heavy" (#986).
- Wrist-camera attention fails on SO-101 (#988).
- When to choose it: over GR00T for arm or mobile-manipulator generalization. Pick GR00T N1.7 for humanoids and license clarity, and SmolVLA for cheap hardware.
- Sources: https://github.com/Physical-Intelligence/openpi/releases ; https://github.com/Physical-Intelligence/openpi/issues ; https://www.pi.website/ ; https://www.pi.website/blog/pi0 ; https://huggingface.co/lerobot/pi05_base ; https://huggingface.co/docs/lerobot/pi05 ; https://huggingface.co/docs/lerobot/libero ; https://en.wikipedia.org/wiki/Physical_Intelligence_Inc.
NVIDIA Isaac GR00T N1.5 / N1.6 / N1.7 (and the N2 preview) — the open humanoid-oriented VLA line
- Status: ACTIVE.
- GTC (Mar 16, 2026): N1.7 announced in early access with commercial licensing.
- Developer forum early-access post: Apr 17, 2026.
- The GitHub README now calls N1.7 "General Availability".
- HF blog (Jul 7, 2026): GR00T 1.7 integrated into LeRobot.
- Issues were still being opened in Aug 2026.
- N1.7 is a BEHAVIOR 2026 baseline.
- Who: NVIDIA's GEAR lab (Jim Fan and Yuke Zhu; prior knowledge) and the Isaac team.
- Links:
- GitHub
NVIDIA/Isaac-GR00T. - HF:
nvidia/GR00T-N1.7-3B,-LIBERO,-DROID,-SimplerEnv-Bridge,-SimplerEnv-Fractal;nvidia/GR00T-N1.6-3B;nvidia/GR00T-N1.5-3B. - White paper: arXiv 2503.14734.
- GitHub
- Stats:
- GitHub: 8.2k stars, 1.5k forks, 213 open issues.
- N1.7-3B: 147,412 downloads last month on the card (45k on the listing), 141 likes.
- N1.6-3B: 29,218 and 91 likes.
- N1.5-3B: 1,402 and 197 likes.
- N1.7-SimplerEnv-Bridge: 532.
- License:
- Code: Apache-2.0.
- Weights:
- N1.5: "Nvidia License (Non-commercial)".
- N1.6: "NVIDIA OneWay Noncommercial License".
- N1.7: NVIDIA Open Model License, "ready for commercial/non-commercial use".
- What it is (N1.7):
- Architecture:
- 3B parameters.
- VLM backbone: Cosmos-Reason2-2B, which uses the Qwen3-VL architecture.
- Action head: flow-matching DiT.
- The state vector grew from 29 to 132 dimensions, with a 40-step action horizon.
- Pretraining: 20K hours of EgoScale human video plus robot demonstrations, using a relative end-effector action space shared across humans and robots.
- Capabilities: task- and subtask-level reasoning, and finger-level dexterous control.
- Embodiments:
- Humanoids: Unitree G1 and AGIBot G1, with whole-body control through the GEAR-SONIC controller.
- Arms: Franka (LIBERO and DROID), WidowX, Google Robot.
- Custom robots via a NEW_EMBODIMENT tag.
- N1.6 uses an Eagle-family VLM (per the model card's citations) and was trained on bimanual, semi-humanoid and humanoid data, both real and synthetic.
- Architecture:
- Compute:
- Inference: one GPU with 16 GB or more (RTX 4090, L40, H100, Jetson AGX Thor or Orin, DGX Spark).
- Model-card latency (4 denoising steps, one camera): 27.9 ms on H100 with TensorRT (≈36 Hz); 216.5 ms on Orin (≈4.6 Hz).
- Fine-tuning: GPUs with 40 GB or more recommended. An open issue (Jul 2026) asks how to fine-tune in bf16 on a 24 GB GPU.
- Adoption:
- NVIDIA names N1.7 adopters (self-reported): AGIBOT, Humanoid, LG Electronics, NEURA Robotics, Noble Machines.
- Isaac Teleop, its data-collection tool, supports an SO-101 leader arm or a VR headset.
- Benchmarks:
- LIBERO via LeRobot (NVIDIA and HF blog): N1.7 Spatial 95%, Object 100%, average 96.5%, versus 87% for GR00T 1.5.
- Third-party tables:
- GR00T-N1 LIBERO 93.9 (94.4 / 97.6 / 93.0 / 90.6), WidowX 45.0.
- GR00T N1.5 WidowX 61.9.
- LingBot-VLA claims to beat N1.6 on GM-100 and RoboTwin 2.0.
- N2 claims (self-reported, Mar 2026): new tasks succeed in new environments "more than twice as often"; "#1 on MolmoSpaces and RoboArena".
- Trade-offs:
- Strengths:
- The only major open line built around humanoids and hands that now has a commercial license.
- NVIDIA's toolchain around it: Isaac Lab 3.0, Cosmos, TensorRT, Jetson Thor.
- Weaknesses:
- Version churn: three releases in about a year.
- Recent issues show confusion: README versus shipped select_layer, padded action dimensions hurting predictions, requests for immutable checkpoint locators, and an install-script package-name mismatch.
- Earlier weights are non-commercial.
- Only about 4.6 Hz on Orin.
- Strengths:
- Sources: https://github.com/NVIDIA/Isaac-GR00T ; https://github.com/NVIDIA/Isaac-GR00T/issues ; https://huggingface.co/nvidia/GR00T-N1.7-3B ; https://huggingface.co/nvidia/GR00T-N1.6-3B ; https://huggingface.co/nvidia/GR00T-N1.5-3B ; https://huggingface.co/nvidia/GR00T-N1.7-SimplerEnv-Bridge ; https://forums.developer.nvidia.com/t/clarification-on-gr00t-license-for-commercial-use/338679 ; https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916 ; https://huggingface.co/blog/nvidia/nvidia-isaac-teleop-and-gr00t17-in-lerobot ; https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world
OpenVLA and OpenVLA-OFT — the 7B academic baseline and its optimized fine-tuning recipe
- Status: OpenVLA is DORMANT; its last README news is from 2025-03-03 and points to OFT. OFT is DORMANT as a repo (paper Feb 2025, revised Apr 2025) but is still the standard baseline.
- Who: Stanford, UC Berkeley and collaborators. OFT is by Moo Jin Kim, Chelsea Finn and Percy Liang.
- Links:
- GitHub
openvla/openvlaandmoojink/openvla-oft. - HF:
openvla/openvla-7b;moojink/openvla-7b-oft-finetuned-libero-{spatial,object,goal,10,spatial-object-goal-10}. - Papers: arXiv 2406.09246 and 2502.19645.
- GitHub
- Stats:
- openvla: 6.8k stars, 820 forks.
openvla-7b: 460,364 downloads last month on the card (169k on the listing), 260 likes. This is the most-downloaded robotics model on HF.- OFT: 1.3k stars, 195 forks. Its LIBERO-spatial checkpoint has 13.9k downloads (listing).
- License: MIT for both. The backbone is Llama-2, so the Llama 2 Community License likely also applies (my note).
- What it is:
- OpenVLA: about 7–8B. DINOv2 plus SigLIP vision encoders with a Llama-2-7B language model, trained on 970K OXE episodes. Actions are discretized tokens (autoregressive).
- OFT changes the fine-tuning recipe: parallel decoding, action chunking, continuous actions and an L1 regression loss. Self-reported results: 26× throughput, LIBERO up from 76.5 to 97.1, and up to 15 points better than π0 and RDT-1B on a real ALOHA robot.
- Compute: OFT inference needs about 16 GB (LIBERO) or about 18 GB (ALOHA). Training uses 1–8 GPUs with 27–80 GB each.
- Benchmarks:
- LIBERO, self-reported:
- OpenVLA: 84.7 / 88.4 / 79.2 / 53.7, average 76.5.
- OFT: 97.6 / 98.4 / 97.9 / 94.5, average 97.1.
- LIBERO-Plus (independent):
- OpenVLA collapses: camera 1.1, robot start 4.1, language 26.8.
- OFT is the most robust model tested: camera 59.7, robot 37.2, language 81.5, light 85.8, background 92.4, noise 76.7, layout 77.1.
- SimplerEnv WidowX for OFT: 63.0 (as cited in the X-VLA table).
- LIBERO, self-reported:
- Trade-offs: OFT is the go-to single-arm fine-tuning baseline and is robust. It is 7B, heavy, has a 2024-era pretraining mix and the Llama-2 license.
- Sources: https://github.com/openvla/openvla ; https://huggingface.co/openvla/openvla-7b ; https://github.com/moojink/openvla-oft ; https://arxiv.org/abs/2502.19645 ; https://arxiv.org/html/2510.13626v1
SmolVLA (Hugging Face) — a 450M community VLA distributed through LeRobot
- Status: ACTIVE through LeRobot.
- lerobot 0.6.1 shipped Aug 3, 2026, after 0.6.0 (Jul 6), 0.5.1 (Apr 7) and 0.5.0 (Mar 9).
- New SmolVLA checkpoints for LIBERO, MetaWorld, RoboTwin and VLABench appeared Mar–Apr 2026.
- I found no SmolVLA successor.
- Who: Hugging Face's LeRobot team. Paper arXiv 2506.01844. The LeRobot library paper is ICLR 2026, arXiv 2602.22818.
- Links: GitHub
huggingface/lerobot; HFlerobot/smolvla_base,lerobot/smolvla_libero,HuggingFaceVLA/smolvla_libero. - Stats:
- lerobot: 27.8k stars, 5.7k forks.
smolvla_base: 159,287 downloads last month on the card (55–71k on the listing), 445 likes, 7,808 fine-tunes, 6 adapters, 13 Spaces.
- License: Apache-2.0.
- What it is:
- Architecture: 450M parameters. SmolVLM2-500M backbone (SigLIP plus SmolLM2) with a ~100M flow-matching action expert.
- Pretraining data: 481 community datasets, 22.9K episodes, 10.6M frames, mostly SO-100 arms at 30 fps (the blog says 487 datasets).
- Asynchronous inference: about 30% faster response and 2× task throughput.
- Training: about 4 hours for 20k steps on one A100.
- The docs say to use at least 50 episodes; "25 episodes… was not enough".
- Benchmarks: SO-100 real-robot success 78.3%, versus 51.7% without pretraining (self-reported). LIBERO average about 87.3% (paper; not re-verified).
- Trade-offs: The cheapest model to fine-tune and run, with the biggest community. It has a lower accuracy ceiling, and its pretraining is skewed toward low-cost arms, not humanoids or hands.
- Sources: https://huggingface.co/lerobot/smolvla_base ; https://huggingface.co/blog/smolvla ; https://huggingface.co/docs/lerobot/smolvla ; https://github.com/huggingface/lerobot ; https://pypi.org/project/lerobot/ ; https://arxiv.org/html/2506.01844
Octo — the first open generalist transformer policy (2024)
- Status: DORMANT. Version 1.5 is from 2024.
- Who: UC Berkeley and collaborators.
- Links: GitHub
octo-models/octo; HFrail-berkeley/octo-base-1.5. - Stats: 1.7k stars, 281 forks. HF: 115 downloads last month, 18 likes.
- License: MIT.
- What it is: Octo-Small (27M) and Octo-Base (93M), trained on 800k OXE trajectories. JAX. Pretraining took 8–14 hours on a TPUv4-128. Runs at 13–17 iterations per second on an RTX 4090.
- Benchmarks: LIBERO 75.1; SimplerEnv WidowX 16.8 (both from the X-VLA comparison table).
- Trade-offs: Tiny and fast but obsolete; superseded by SmolVLA and X-VLA.
- Sources: https://github.com/octo-models/octo ; https://huggingface.co/rail-berkeley/octo-base-1.5
B. Open VLAs from 2025–2026 (industry and academia)
X-VLA (Tsinghua AIR + Shanghai AI Lab) — 0.9B cross-embodiment VLA with soft prompts
- Status: SLOWING. Accepted at ICLR 2026; HF collection updated Mar 2, 2026; maintained in LeRobot. Won the IROS 2025 AgiBot World Challenge.
- Links: GitHub
2toinf/X-VLA; HF2toINF/X-VLA-Pt,2toINF/X-VLA-Libero,lerobot/xvla-libero; arXiv 2510.10274. - Stats: 728 stars, 71 forks. X-VLA-Pt: 10,471 downloads last month, 14 likes.
- License: Apache-2.0.
- What it is: 0.9B. A Florence-2 encoder plus a transformer with learnable "soft prompts" per embodiment. Flow matching over a 6D end-effector action space. Pretrained on 290K episodes from DROID, RoboMind and AgiBot, covering 7 platforms and 5 arm types.
- Benchmarks (self-reported):
- LIBERO 98.1.
- SimplerEnv Google Robot: visual matching (VM) 83.5, variant aggregation (VA) 76.4. WidowX 95.8.
- CALVIN ABC→D 4.43. VLABench 51.1. RoboTwin 2.0 70%. Cloth folding 100%.
- RoboChallenge Table30: 21.33% (aggregator).
- Trade-offs: The best sim accuracy per parameter. Weak on real-robot RoboChallenge, about half of π0.5.
- Sources: https://github.com/2toinf/X-VLA ; https://huggingface.co/2toINF/X-VLA-Pt ; https://arxiv.org/html/2510.10274
RDT-1B and RDT2 (Tsinghua) — a diffusion foundation model, then a UMI-data-scaled VLA
- Status: RDT-1B is DORMANT (last news Apr 2025). RDT2 is ACTIVE: models Sept 2025, paper Feb 2026, accepted at ICML 2026.
- Links:
- GitHub
thu-ml/RoboticsDiffusionTransformerandthu-ml/RDT2. - HF
robotics-diffusion-transformer/rdt-1bandrobotics-diffusion-transformer/RDT2-VQ(plus RDT2-FM). - Papers: arXiv 2410.07864 and 2602.03310.
- GitHub
- Stats:
- RDT-1B: 1.8k stars, 163 forks; HF 405 downloads last month, 106 likes.
- RDT2: 805 stars, 58 forks; RDT2-VQ 371 on the card (13.7k on the listing), 22 likes.
- License: RDT-1B MIT. RDT2 Apache-2.0.
- What it is:
- RDT-1B: a 1B diffusion transformer pretrained on 1M+ multi-robot episodes and fine-tuned on 6K+ bimanual ALOHA episodes. Predicts 64-step chunks at about 25 Hz control.
- RDT2: a Qwen2.5-VL-7B backbone in two variants. RDT2-VQ (8B) is autoregressive with residual-VQ action tokens. RDT2-FM uses a flow-matching expert for lower latency.
- RDT2 data: 10,000+ hours of human manipulation collected with an improved handheld UMI gripper in 100+ indoor scenes.
- RDT2 claims zero-shot deployment on unseen embodiments, tested on bimanual UR5e and Franka FR3.
- Compute (RDT2):
- Inference: RTX 4090 (~16 GB).
- Fine-tuning: FM on a 4090; VQ with LoRA on an A100-40GB; VQ full fine-tuning needs ≥80 GB.
- Benchmarks: RDT-1B scores 15% on RoboChallenge (aggregator). RDT2's zero-shot claims are self-reported.
- Trade-offs: RDT2 is the clearest open evidence that wearable or handheld human data capture scales into policies. That makes it directly relevant to a data business. It is limited to gripper tasks.
- Sources: https://github.com/thu-ml/RDT2 ; https://github.com/thu-ml/RoboticsDiffusionTransformer ; https://huggingface.co/robotics-diffusion-transformer/RDT2-VQ ; https://huggingface.co/robotics-diffusion-transformer/rdt-1b
WALL-OSS (X Square Robot) — a 4B embodied foundation model
- Status: ACTIVE.
- WALL-OSS-0.5 released May 2026 ("Gradient-Bridged Pretraining").
- Wall-X 1.1.0 (training and inference stack) released June 2026.
- WALL-WM world-action model announced May 2026.
- Original release Sept 2025.
- Who: X Square Robot, Shenzhen. Press headlines report a $140M Series A++ around June 2026 (headline only).
- Links: GitHub
X-Square-Robot/wall-x; HFx-square-robot/wall-oss-flow,wall-oss-fast, WALL-OSS-0.5. - Stats: 1.3k stars, 99 forks. wall-oss-flow: 908 downloads last month, 34 likes.
- License: Apache-2.0 (repo).
- What it is: 4B, built on Qwen2.5-VL with mixture-of-experts action layers. Comes in a flow-matching variant and a FAST-token variant. Data hours are not disclosed.
- Benchmarks: RoboChallenge 35.33% (aggregator).
- Trade-offs: Permissive license and mid-tier real-robot results; data documentation is thin.
- Sources: https://github.com/X-Square-Robot/wall-x ; https://huggingface.co/x-square-robot/wall-oss-flow
AgiBot GO-1 (open) and GO-2 (closed)
- Status:
- GO-1 is SLOWING: open-sourced Sept 19, 2025; the paper is listed as IEEE TRO 2026.
- GO-2 was announced Apr 9, 2026 and is not open; it does not appear in the AgiBot-World repo.
- Who: AgiBot (Shanghai) with OpenDriveLab (HKU).
- Links: GitHub
OpenDriveLab/AgiBot-World; HFagibot-world/GO-1,agibot-world/GO-1-Air. - Stats: 3.2k stars, 218 forks. GO-1: 149 downloads last month, 20 likes.
- License: CC BY-NC-SA 4.0 for code, data and weights, so non-commercial.
- What it is:
- GO-1: 3B on an InternVL2.5-2B backbone. A latent planner plus a diffusion action expert; GO-1 Air drops the planner. Full fine-tuning needs about 70 GB.
- Data: AgiBot World Beta has 1,003,672 trajectories (~43.8T tokens) from 100 robots.
- GO-2 (self-reported): "action chain-of-thought" plus an asynchronous dual system, trained on "tens of thousands of hours".
- GO-2 benchmarks (self-reported): LIBERO 98.5; LIBERO-Plus 86.6 zero-shot; VLABench 47.4; Genie Sim 3.0 sim-only training transferred to real at 82.9%.
- Sources: https://github.com/OpenDriveLab/AgiBot-World ; https://huggingface.co/agibot-world/GO-1 ; https://www.therobotreport.com/agibot-releases-go-2-foundation-model-embodied-ai/
InternVLA-M1 and InternVLA-A1 (Shanghai AI Lab)
- Status: M1 is SLOWING (Oct 2025). A1 is SLOWING: paper Jan 5, 2026, revised Feb 13; HF updated Mar 27, 2026.
- Links:
- GitHub
InternRobotics/InternVLA-M1and the InternRobotics "InternVLA-A-series" repo (exact repo name not captured). - HF
InternRobotics/InternVLA-M1andInternRobotics/InternVLA-A1-3B. - Papers: arXiv 2510.13778 and 2601.02456.
- GitHub
- Stats:
- M1: 432 stars.
- A-series repo: 479 stars.
- A1-3B: 81 downloads last month, 40 likes.
- License: M1 is MIT. A1 weights are CC BY-NC-SA 4.0 (non-commercial).
- What it is:
- M1: about 4.1B. Qwen2.5-VL-3B plus DINOv2, with an 86M diffusion action head. Pretrained on spatial grounding before action training ("spatially guided").
- A1: 2B or 3B, built on InternVL3 or Qwen3-VL. A mixture of transformers with three experts: understanding, generation (predicting future frames) and action.
- A1 data: 692M+ frames mixing real-robot data, synthetic simulation data (InternData-A1) and human video.
- Benchmarks (self-reported):
- M1: LIBERO 95.9 (98.0 / 99.0 / 93.8 / 92.6); SimplerEnv Google VM 80.7 / VA 76.0; WidowX 71.7.
- A1: +4.4% over π0.5 on static tasks, +26.7% on dynamic tasks, +2.6% on RoboTwin 2.0. The HF card gives RoboTwin 2.0 Easy 89.40 and Hard 89.64.
- An "A1" entry scores 29% on RoboChallenge (aggregator; that it is InternVLA-A1 is my unverified guess).
- Trade-offs: A heavy user of synthetic data, which is a good signal for simulation vendors. A1 is non-commercial.
- Sources: https://arxiv.org/abs/2601.02456 ; https://arxiv.org/html/2510.13778 ; https://github.com/InternRobotics/InternVLA-M1 ; https://github.com/InternRobotics ; https://huggingface.co/InternRobotics/InternVLA-A1-3B
MolmoAct and MolmoAct2 (AI2) — the most fully open 2026 release
- Status: ACTIVE.
- MolmoAct2 released May 4–5, 2026.
- Full experiment code June 2026; LeRobot fine-tuning fixes Aug 2026.
- MolmoAct v1 dates from Aug 12, 2025.
- Who: Allen Institute for AI. MolmoAct2 is led by Haoquan Fang and Jiafei Duan with 20+ contributors.
- Links:
- GitHub
allenai/molmoactandallenai/molmoact2. - HF
allenai/MolmoAct2,-Think,-Pretrain,-LIBERO,-DROID,-BimanualYAM,-SO100_101. - Paper: arXiv 2605.02881.
- GitHub
- Stats:
- molmoact: 383 stars.
- molmoact2: 782 stars, 63 forks.
MolmoAct2: 19,785 downloads last month, 24 likes.MolmoAct2-LIBERO: 4–21k (listing).
- License: Apache-2.0. Weights, training code and the complete training data are released.
- What it is:
- 5B, built on Molmo2-ER, an embodied-reasoning VLM trained on 3.3M samples. AI2 claims it beats GPT-4 and Gemini Robotics-ER 1.5 on 13 embodied-reasoning benchmarks.
- A flow-matching action expert attached through per-layer KV conditioning.
- Also released: OpenFAST, an open-weight, open-data version of the FAST action tokenizer, and MolmoThink, adaptive-depth reasoning.
- Data: 720 h of bimanual YAM teleoperation and re-annotated robot data (unique labels grew from 71k to about 146k). Five embodiments; runs on Intel XPU as well.
- AI2 says it is "up to 37× faster" than its predecessor.
- Benchmarks: Claims to outperform π0.5 across 7 sim and real benchmarks (self-reported; numbers not captured). Known limitations: the gripper blocking the camera view, and fine-grained tasks.
- Trade-offs: The best base to build on if openness and provenance matter (US nonprofit, Apache license, data included). No independent numbers yet.
- Sources: https://huggingface.co/allenai/MolmoAct2 ; https://huggingface.co/papers/2605.02881 ; https://github.com/allenai/molmoact2 ; https://github.com/allenai/molmoact ; https://siliconangle.com/2026/05/05/ai2-releases-molmoact-2-enhancing-robot-intelligence-real-world/
Xiaomi-Robotics-0 — 4.7B VLA with top self-reported simulation scores
- Status: ACTIVE. Report Feb 13, 2026; post-training code Apr 27, 2026; HF checkpoint updated mid-Sept 2026.
- Links:
- GitHub
XiaomiRobotics/Xiaomi-Robotics-0. - HF
XiaomiRobotics/Xiaomi-Robotics-0-{Pretrain, LIBERO, Calvin-ABCD_D, Calvin-ABC_D, SimplerEnv-Google-Robot, SimplerEnv-WidowX}. - Paper: arXiv 2602.12684.
- GitHub
- Stats: 641 stars, 70 forks. LIBERO checkpoint: 3.34k downloads (listing).
- License: Apache-2.0.
- What it is: 4.7B. Qwen3-VL-4B plus a 16-layer DiT flow-matching head. Data: about 200M robot timesteps plus 80M+ vision-language samples, plus in-house data (338 h Lego disassembly, 400 h towel folding). 80 ms inference on an RTX 4090.
- Benchmarks (self-reported): LIBERO 98.8 / 100 / 98.8 / 97.2, average 98.7. CALVIN ABCD→D 4.80, ABC→D 4.75. SimplerEnv Google VM 85.5 / VA 74.7; WidowX 79.2.
- Sources: https://arxiv.org/html/2602.12684v1 ; https://github.com/XiaomiRobotics/Xiaomi-Robotics-0
Spirit v1.5 (Spirit AI, Beijing) — #1 on RoboChallenge when released
- Status: ACTIVE. Released Jan 2026 (press release Jan 12); fine-tuning code Apr 2026.
- Links: GitHub
Spirit-AI-Team/spirit-v1.5; HFSpirit-AI-robotics/Spirit-v1.5. - Stats: 624 stars, 36 forks. HF: 131 downloads last month, 31 likes.
- License: MIT on GitHub, Apache-2.0 on HF.
- What it is: 5B. Qwen3-VL-4B plus a DiT action head. Trained "largely on open-ended, goal-driven diverse data" rather than scripted demonstrations. Tested on A100-80GB; 8×A100 recommended for training.
- Benchmarks: Ranked #1 on RoboChallenge Table30 as of Jan 11, 2026 (self-reported). The aggregator snapshot shows 51%.
- Why it matters for data vendors: It is evidence that unscripted, goal-driven data collection pays off.
- Sources: https://github.com/Spirit-AI-Team/spirit-v1.5 ; https://huggingface.co/Spirit-AI-robotics/Spirit-v1.5 ; https://finance.yahoo.com/news/robochallenges-top-ranked-embodied-ai-064100221.html
GigaBrain (GigaAI) — GigaBrain-0 and 0.7
- Status: ACTIVE. GigaBrain-0.7 published Aug 16, 2026 (arXiv 2608.15875).
- Links: HF
open-gigaai/GigaBrain-0.7-3.5B-Baseandopen-gigaai/Giga-World-Policy-0.5(updated Jul 30, 2026). - Stats: 0.7-Base: 6,005 downloads last month, 12 likes.
- License: Apache-2.0.
- What it is: About 3.5–4B. A "three-system architecture" that combines understanding, prediction and action. Trained on 37,000+ hours of mixed embodied data.
- Benchmarks: "GigaBrain" scores 51.67% on RoboChallenge (aggregator; version unspecified).
- Sources: https://huggingface.co/open-gigaai/GigaBrain-0.7-3.5B-Base ; https://www.sota2.com/research/sota/robotic-manipulation-on-robochallenge-table30
Dexmal DM0 / DM0.5 (OpenDM) — from the co-organizer of RoboChallenge
- Status: ACTIVE.
- DM0.5 released Jul 9, 2026.
- OpenDM policy added to XPolicyLab Sept 23, 2026.
- Claimed #1 on all four RoboColiseum leaderboards Sept 14, 2026.
- Who: Dexmal (Beijing). It co-authored RoboChallenge with Hugging Face, which is a conflict of interest for its RoboChallenge scores.
- Links:
- GitHub
Dexmal/opendm,Dexmal/dexbotic,Dexmal/realtime-vla. - HF DM05 checkpoints: base, -libero, -robotwin2, -SO101-Pick-Cube, -Vla-Arena, -MEM-Robodojo-Sim.
- GitHub
- Stats:
- opendm: 2.2k stars and 184 forks on the repo page; the org page showed 304, a discrepancy.
- dexbotic (VLA toolbox): 1,389 stars.
- realtime-vla: 604 stars ("30Hz frame rate and 480Hz trajectory").
- License: opendm Apache-2.0; dexbotic and realtime-vla MIT.
- What it is: Parameter count and backbone are not disclosed. Inference fits on one GPU (RTX 4090 up to H100); 8 GPUs recommended for training.
- Benchmarks (self-reported):
- LIBERO 99.0.
- RoboTwin 2.0: 93.6 clean, 93.3 randomized.
- RoboChallenge Table30v2: score 54.42, 43.0% success.
- RoboDojo-Sim: score 24.90, 19.34% success.
- DM0 scores 62% on Table30 in the aggregator snapshot, the top entry.
- Sources: https://github.com/Dexmal/opendm ; https://github.com/Dexmal
LingBot-VLA (Robbyant, Ant Group's embodied AI unit)
- Status: ACTIVE. v1 released Jan 27, 2026;
lingbot-vla-v2-6b-robotwinposted Jul 24, 2026. The org also released LingBot-World-V2 and LingBot-Video models in Jul 2026. - Links: GitHub
Robbyant/lingbot-vla; HFrobbyant/lingbot-vla-v2-6b-robotwin. - License: Apache-2.0.
- What it is: 4B on Qwen2.5-VL-3B, with and without depth input. Trained on 20,000 h of real-world data from 9 dual-arm robot configurations. Claims 1.5–2.8× faster training than other VLA codebases.
- Benchmarks (self-reported): RoboTwin 2.0 88.56 clean and 86.68 randomized (with depth). GM-100 average success across 3 platforms is only 15.74% (no depth) or 17.30% (depth), but the authors say this beats π0.5 and GR00T N1.6.
- Sources: https://github.com/Robbyant/lingbot-vla ; https://huggingface.co/robbyant
Galaxea G0 → G0.5 (Galaxea AI) — VLAs for wheeled mobile manipulators
- Status: ACTIVE.
- G0.5 released Jun 2026.
- G0Tiny (250M, for the R1 Pro's onboard Orin) Feb 2026.
- G0Plus (3B, pretrained on 2k+ h of teleoperation) Jan 2026.
- G0 Sept 2025.
- Links: GitHub
OpenGalaxea/GalaxeaVLA; HFOpenGalaxea/G05(g05-base, -droid, -so101, -libero, -robotwin20) andOpenGalaxea/G0-VLA. - Stats: 759 stars, 57 forks.
- License: Content from before Jan 2026 is Apache-2.0. G0.5 and later use a non-commercial community license.
- What it is:
- G0.5: about 2B on Qwen3.5-2B. One autoregressive stream of reasoning tokens and action tokens, using a 27-dimensional action codec that includes the lower body, plus multi-frame visual memory.
- Data: the Galaxea Open-World Dataset (500+ h of mobile manipulation) and 14 embodiments.
- Compute: inference >8 GB; fine-tuning >70 GB.
- Benchmarks (self-reported): LIBERO 98.9; SimplerEnv Bridge 87.3; RoboTwin 2.0 93.3; BEHAVIOR-1K 0.3136; DROID 82.5% zero-shot.
- Sources: https://github.com/OpenGalaxea/GalaxeaVLA
EO-1 (Shanghai AI Lab / EO-Robotics) — interleaved reasoning and action
- Status: SLOWING. The repo was archived Nov 12, 2025 and development moved to SHAILAB-IPEC/EO1 (not fetched). EO-1 is supported in LeRobot.
- Links: GitHub
EO-Robotics/EO1; HFIPEC-COMMUNITY/EO-1-3B; arXiv 2508.21112. - Stats: 291 stars. HF: 55 downloads last month, 14 likes.
- License: MIT.
- What it is: 3B (HF says 4B) on Qwen2.5-VL-3B. One decoder that does both autoregressive text and flow-matching action generation. Trained on EO-Data1.5M; inference needs about 6.5 GB.
- Benchmarks (self-reported): LIBERO 98.2 (99.7 / 99.8 / 99.2 / 94.8); SimplerEnv WidowX 72.7.
- Sources: https://github.com/EO-Robotics/EO1 ; https://huggingface.co/IPEC-COMMUNITY/EO-1-3B
Unitree UnifoLM-VLA-0 — an open VLA aimed at the G1 humanoid
- Status: SLOWING. Latest release Jan 29, 2026.
- Links: GitHub
unitreerobotics/unifolm-vla; HFunitreerobotics/UnifoLM-VLM-Base,UnifoLM-VLA-Base,UnifoLM-VLA-Libero. - Stats: 546 stars, 57 forks.
- License: Not captured.
- What it is: Based on Qwen2.5-VL and borrowing from GR00T, OpenVLA-OFT and InternVLA-M1. Trained on 12 open datasets plus Unitree G1 manipulation data; one policy covers 12 real task categories.
- Sources: https://github.com/unitreerobotics/unifolm-vla
VLA-JEPA — a VLA with a world-model loss used only during training
- Status: ACTIVE. Paper arXiv 2602.10098 (Feb 2026); ported to LeRobot.
- Links: HF
lerobot/VLA-JEPA-Pretrain,lerobot/VLA-JEPA-LIBERO,lerobot/VLA-JEPA-SimplerEnv. - Stats: Pretrain checkpoint: 2.78k downloads (listing).
- License: Apache-2.0.
- What it is: Qwen3-VL-2B plus a DiT-B flow-matching head. A V-JEPA 2 action-conditioned video predictor adds a self-supervised loss during training only, so it costs nothing at inference. Pretrained on DROID.
- Benchmarks: LeRobot run on LIBERO: 95.0 / 100 / 98.0 / 93.0, overall 96.5 (100 episodes per suite).
- Sources: https://huggingface.co/docs/lerobot/vla_jepa
CogACT (Microsoft Research Asia) — VLM plus diffusion action transformer
- Status: DORMANT. Initial release Dec 1, 2024; last update Dec 23, 2024.
- Links: GitHub
microsoft/CogACT; HF CogACT/CogACT-{Small, Base, Large}. - Stats: 432 stars.
- License: MIT for weights, code and data.
- What it is: Diffusion transformer action heads (S, B or L sizes) over 16-step chunks. About 181 ms per inference on an A6000 (about 5.5 Hz with its adaptive action ensemble).
- Benchmarks: SimplerEnv Google VM 74.8 / VA 61.3; WidowX 51.3 (InternVLA-M1 table). RoboChallenge 11.7% (organizer-run).
- Sources: https://github.com/microsoft/CogACT ; https://arxiv.org/html/2510.13778 ; https://arxiv.org/html/2510.17950v1
SpatialVLA — 4B VLA with 3D-aware action tokens
- Status: DORMANT. Last update Mar 2025.
- Links: GitHub
SpatialVLA/SpatialVLA; HFIPEC-COMMUNITY/spatialvla-4b-224-pt. - Stats: 727 stars.
- License: MIT. It is built on PaliGemma2-3B, so Gemma terms likely apply to the weights.
- What it is: Trained on 1.1M real episodes (OXE and RH20T) using 64 A100s for about 10 days. Inference needs about 8.5 GB.
- Benchmarks (self-reported): SimplerEnv Google Robot 71.9% zero-shot (whether VM or VA is not specified); WidowX 34.4% zero-shot; LIBERO 78.1. The InternVLA-M1 table lists 75.1 VM / 70.7 VA.
- Sources: https://github.com/SpatialVLA/SpatialVLA
UniVLA (OpenDriveLab) — task-centric latent actions learned from videos
- Status: DORMANT. v1.0 code released May 2025.
- Links: GitHub
OpenDriveLab/UniVLA; HF univla-7b and variants. - Stats: 1.1k stars.
- License: Apache-2.0.
- What it is: 7B. Learns latent actions with a VQ-VAE from OXE plus an Ego4D subset of human video. Pretraining took about 960 A100 GPU-hours. Adapts with a 12M decoder plus LoRA.
- Benchmarks: LIBERO 96.5 / 96.8 / 95.6 / 92.0 (average about 95.2); WidowX 47.9 (self-reported).
- Sources: https://github.com/OpenDriveLab/UniVLA
Being-H0 (BeingBeyond) — VLA pretrained on human hand-motion video
- Status: ACTIVE as a series: accepted at ICML 2026 (May 1, 2026); code released Aug 2025. The README refers to newer Being-H0.5 and H0.7 (details not verified).
- Links: GitHub
BeingBeyond/Being-H0. - Stats: 59 stars.
- License: MIT.
- What it is: 1B, 8B and 14B models plus a motion tokenizer. Models explicit hand motion with MANO hand models for dexterous manipulation.
- Why it matters: One of the few open models aimed at dexterous hands, and directly relevant to glove and hand-tracking data.
- Sources: https://github.com/BeingBeyond/Being-H0
RoboBrain 2.0 / 2.5 (BAAI) — an embodied planning model; it does not output actions
- Status: SLOWING. RoboBrain 2.5 (4B and 8B) Jan 2026; Robo-Dopamine (temporal value estimation) Dec 2025.
- Links: GitHub
FlagOpen/RoboBrain2.0; HFBAAI/RoboBrain2.5-8B-NV,BAAI/RoboBrain2.5-4B,BAAI/RoboBrain2.0-7B,-32B. - Stats: 1.1k stars.
- License: Apache-2.0.
- What it is: 3D spatial reasoning, dense temporal value or progress estimation, and planning. It feeds downstream controllers.
- Sources: https://github.com/FlagOpen/RoboBrain2.0
Magma (Microsoft) — an agent model spanning UI navigation and robotics
- Status: DORMANT. Last news Apr 29, 2025.
- Links: GitHub
microsoft/Magma; HF Magma-8B. - Stats: 1.9k stars.
- License: MIT.
- What it is: 8B on Llama-3-8B, pretrained with Set-of-Mark and Trace-of-Mark visual annotations. CVPR 2025.
- Sources: https://github.com/microsoft/Magma
RynnVLA-002 (Alibaba DAMO) — autoregressive action world model (formerly WorldVLA)
- Status: SLOWING. Released Nov 10, 2025.
- Links: GitHub
alibaba-damo-academy/RynnVLA-002; HF Alibaba-DAMO-Academy/RynnVLA-002. - Stats: 1.1k stars.
- License: Apache-2.0.
- What it is: Combines action prediction (VLA) with next-image prediction (world model) in one model.
- Benchmarks: LIBERO suites 94.2–99.8 (self-reported). Its predecessor WorldVLA scores 79.1 on LIBERO and collapses to 0.3 under camera changes in LIBERO-Plus.
- Sources: https://github.com/alibaba-damo-academy/RynnVLA-002 ; https://arxiv.org/html/2510.13626v1
VLA-Adapter (OpenHelix) — a tiny-backbone VLA
- Status: ACTIVE / SLOWING. ALOHA support added Mar 2026; paper Sept 2025.
- Links: GitHub
OpenHelix-Team/VLA-Adapter. - Stats: 2.3k stars.
- License: MIT.
- What it is: Qwen2.5-0.5B backbone. Trains in about 8–10 h on one 40 GB GPU.
- Benchmarks (Pro version, self-reported): LIBERO 99.6 / 99.6 / 98.2 / 96.4, average 98.5; CALVIN ABC→D 4.50.
- Sources: https://github.com/OpenHelix-Team/VLA-Adapter
NORA (SUTD declare-lab)
- Status: SLOWING. NORA-1.5 has been released (date not verified).
- Links: GitHub
declare-lab/nora; HFdeclare-lab/nora-long(3.83k downloads, listing). - Stats: 222 stars.
- What it is: Qwen2.5-VL-3B with the FAST+ action tokenizer.
- Benchmarks: LIBERO 87.9; LIBERO-Plus camera-change score 4.0 (independent).
- Sources: https://github.com/declare-lab/nora ; https://arxiv.org/html/2510.13626v1
FLOWER (KIT) — a ~1B VLA that is cheap to train
- Status: DORMANT / SLOWING (2025).
- Links: GitHub
intuitive-robots/flower_vla_calvin. - Stats: 94 stars.
- License: MIT.
- What it is: Florence-2 with a rectified-flow action head. About 200 GPU-hours of pretraining; <3 GB at inference.
- Benchmarks (self-reported): CALVIN ABC→D 4.54 and ABCD→D 4.67; LIBERO 97.2 / 99.3 / 96.9 / 94.5.
- Sources: https://github.com/intuitive-robots/flower_vla_calvin
Dita — a diffusion-transformer generalist policy
- Status: DORMANT (ICCV 2025).
- Links: GitHub
RoboDita/Dita. - Stats: 172 stars.
- License: Apache-2.0.
- Benchmarks: LIBERO 82.4 (self-reported).
- Sources: https://github.com/RoboDita/Dita
Classic single-task policies (still the default small baselines)
These four are all MIT-licensed.
| Policy | GitHub repo | Stars | Status | Notes |
|---|---|---|---|---|
| Diffusion Policy | real-stanford/diffusion_policy | 4.6k | Reference code dormant; the method lives on in LeRobot and elsewhere | RSS 2023 |
| ACT | tonyzhaozh/act | 2.2k | Reference code dormant | LeRobot calls it "recommended first policy" |
| DP3 (3D Diffusion Policy) | YanjieZe/3D-Diffusion-Policy | 1.4k | SLOWING; extensions through Oct 2025 | Point-cloud input; trains in about 3 h on an A40 with ~10 GB; includes Franka + Allegro hand code |
| HPT | liruiw/HPT | 542 | DORMANT | NeurIPS 2024; 3.1M–227M parameters |
- Sources: https://github.com/real-stanford/diffusion_policy ; https://github.com/tonyzhaozh/act ; https://github.com/YanjieZe/3D-Diffusion-Policy ; https://github.com/liruiw/HPT
C. World models and world-action models
NVIDIA Cosmos (Predict 2.5, Transfer 2.5, Reason 2, Cosmos 3 including Nano-Policy)
- Status: ACTIVE. Cosmos 3 checkpoints on HF were updated within hours of my fetch; Cosmos3-Nano-Policy-DROID was released May 31, 2026; cosmos-rl was updated Sept 24, 2026.
- Who: NVIDIA.
- Links:
- GitHub
nvidia-cosmos/{cosmos-predict2.5, cosmos-transfer2.5, cosmos-reason2, cosmos-reason1, cosmos-predict1, cosmos-transfer1, cosmos-cookbook, cosmos-rl}. - HF
nvidia/Cosmos-Predict2.5-2B,nvidia/Cosmos3-Edge,nvidia/Cosmos3-Nano,nvidia/Cosmos3-Super,nvidia/Cosmos3-Super-Image2Video,nvidia/Cosmos3-Nano-Policy-DROID.
- GitHub
- Stats:
- GitHub stars: predict2.5 1,373; transfer2.5 743; reason2 452; reason1 962; transfer1 827; predict1 474; cookbook 477; rl 474.
- HF: Predict2.5-2B 7,631 downloads last month, 167 likes.
- Cosmos 3 listing (30-day): Edge 950k, Nano 282k, Super 121k, Super-Image2Video 74.8k. Nano-Policy-DROID 2,472 on the card (931 on the listing), 34 likes.
- License:
- Code: Apache-2.0.
- Predict2.5: NVIDIA Open Model License, "Models are commercially usable".
- Cosmos3-Nano-Policy-DROID: "OpenMDW1.1", "ready for commercial and non-commercial use".
- What it is:
- Predict2.5 (2.1B, Oct 6, 2025): generates video from text, image or video. Has action-conditioned robot variants (256p, 4 FPS) and policy variants trained on LIBERO and RoboCasa.
- Cosmos 3 (announced Mar 16, 2026): an "omni-model" pairing a reasoner with a generator.
- Edge: 2B + 2B, aimed at real-time robot policy.
- Nano: 8B + 8B.
- Super: 32B + 32B.
- Cosmos3-Nano-Policy-DROID: 16B, a mixture of autoregressive and diffusion transformers. Trained on 1.3B data points from 393 datasets, including OpenImage, YouTube video, UMI and synthetic data. Claims "ranked #1" on RoboArena (self-reported).
- Reason2 is the VLM backbone of GR00T N1.7.
- Adoption: NVIDIA names Cosmos users FieldAI, Skild AI, World Labs, Generalist AI, CMR Surgical and J&J MedTech (self-reported).
- Trade-offs: The best-resourced open world-model family, under commercial licenses. The policy variants are new and need 16B+ parameters.
- Sources: https://github.com/nvidia-cosmos ; https://huggingface.co/nvidia/Cosmos-Predict2.5-2B ; https://huggingface.co/models?search=cosmos3 ; https://huggingface.co/nvidia/Cosmos3-Nano-Policy-DROID ; https://huggingface.co/nvidia?search_models=cosmos ; https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world
DreamZero (NVIDIA GEAR) and GR00T N2
- Status: ACTIVE. DreamZero paper arXiv 2602.15922 (Feb 2026); DROID checkpoint posted Mar 15, 2026. GR00T N2 was previewed Mar 16, 2026 and is planned for end of 2026.
- Links: HF
GEAR-Dreams/DreamZero-DROID. - Stats: 1,423 downloads last month, 34 likes.
- License: CC-BY-NC-4.0 (non-commercial).
- What it is: A 14B "World Action Model" built on the Wan2.1-I2V-14B-480P video-diffusion model. It predicts future frames and actions jointly, and runs closed-loop at about 7 Hz with its "DreamZero-Flash" speed-ups. Self-reported: more than 2× better generalization to new tasks and environments than state-of-the-art VLAs.
- Sources: https://huggingface.co/GEAR-Dreams/DreamZero-DROID ; https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world
Genie Envisioner (AgiBot) — video world model, action decoder and neural simulator
- Status: ACTIVE. GE-Act v2.0 released Sept 10, 2026; GE-Sim v2.0 May 28, 2026; GE-Act CALVIN weights Dec 18, 2025; GE-Base weights Aug 14, 2025.
- Links: GitHub
AgibotTech/Genie-Envisioner. - Stats: 586 stars.
- License: CC BY-NC-SA 4.0 (some directories Apache-2.0).
- What it is:
- GE-Base: a video world model on LTX-Video 2B.
- GE-Act: an action decoder.
- GE-Sim: a neural simulator; v1 was based on Cosmos2.
- Trained on AgiBot World; post-training uses LeRobot-format data.
- Sources: https://github.com/AgibotTech/Genie-Envisioner
Unitree UnifoLM-WMA-0
- Status: SLOWING. Training code and weights Sept 15, 2025; deployment code Sept 22, 2025.
- Links: GitHub
unitreerobotics/unifolm-world-model-action. - Stats: 1.1k stars, 140 forks.
- What it is: A DynamiCrafter video model with a Diffusion Policy action head. It runs in two modes: as a simulation engine that generates synthetic data, and as a policy enhancer. The Base version is trained on Open-X; the Dual version on 5 Unitree datasets (Z1 arm and G1 humanoid).
- Sources: https://github.com/unitreerobotics/unifolm-world-model-action
Meta V-JEPA 2 / V-JEPA 2-AC / V-JEPA 2.1
- Status: ACTIVE / SLOWING. V-JEPA 2.1 released Mar 16, 2026; V-JEPA 2 Jun 25, 2025.
- Links: GitHub
facebookresearch/vjepa2. - Stats: 4.7k stars, 580 forks.
- License: Code is MIT and Apache-2.0.
- What it is:
- Self-supervised video encoders from 80M (ViT-B) to 2B (ViT-G).
- V-JEPA 2-AC adds an action-conditioned predictor post-trained on 62 h of DROID data.
- Zero-shot planning on a Franka arm: reach 100%; grasp a cup 60% and a box 20%; pick-and-place a cup 80% and a box 50% (self-reported).
- Planning is slow, about 16 s per action (paper; not re-verified).
- Sources: https://github.com/facebookresearch/vjepa2
OpenWAM-Alpha
- Status: ACTIVE (arXiv 2609.07398, Sept 2026).
- Links: HF
OpenWAM/OpenWAM-Alpha-Pretrain-Foundation-Model. - Stats: 343 downloads last month.
- License: Apache-2.0.
- What it is: Built on the Wan2.2-TI2V-5B video model. Trained on 518.5M frames (about 6,400 h): 70% robot data (40% real, 30% simulated) and 30% egocentric human data.
- Why it matters: A rare public breakdown of a pretraining data mix.
- Sources: https://huggingface.co/OpenWAM/OpenWAM-Alpha-Pretrain-Foundation-Model
TRI Large Behavior Models — not open
- Status: No weight or code release is mentioned on the project page. The paper appears in Science Robotics (2026).
- What it is: A diffusion transformer that predicts 1.6 s action chunks.
- Data: About 1,700 h: 468 h of internal bimanual teleoperation, 45 h of simulated teleoperation, 32 h of UMI and about 1,150 h of curated OXE.
- Evaluation: 1,800 real rollouts and 47k+ simulated rollouts, run blind A/B with sequential hypothesis testing. Fine-tuning from the pretrained model needs 3–5× less data on hard tasks.
- Sources: https://toyotaresearchinstitute.github.io/lbm1/
D. Closed comparators (brief)
- Google DeepMind Gemini Robotics 2 (Jul 30, 2026): three models.
- Gemini Robotics 2: a VLA for whole-body control and dexterity. Early-access partners only.
- Gemini Robotics ER 2: embodied reasoning, based on Gemini 3.5 Flash. Public preview via the Gemini API.
- On-Device 2: Trusted Testers only.
- Results:
- Five-finger hand tasks succeed 32–92% of the time.
- Whole-body pickup on an Apptronik Apollo 2 succeeds 45.7–76.3%, depending on where the object is.
- On-Device 2 went from 6.7% to 53.3% on an unseen SO101 arm with fewer than 200 examples.
- Only the ASIMOV-Agentic safety benchmark (CC-BY-4.0) and sample code are open.
- Earlier: Gemini Robotics 1.5 (Sept 2025) and On-Device (June 2025) (prior knowledge).
- Generalist AI:
- Model releases (closed):
- GEN-0: Nov 4, 2025.
- GEN-1: Apr 2, 2026.
- GEN-1.5 ("Embodied Foundation Models are One-Shot Learners"): Aug 19, 2026.
- The GEN-0 post claimed about 270k hours of real-world manipulation data (not re-verified).
- Model releases (closed):
- Physical Intelligence π*0.6 (Nov 2025, reinforcement learning from experience) and MEM (Mar 2026) are closed.
- Figure Helix (Feb 2025): a 7B "System 2" model plus an 80M "System 1" running at 200 Hz, trained on about 500 h of teleoperation. 1X has Redwood AI and a world model used for evaluation. Skild Brain launched July 2025, and Skild is named as a Cosmos user. Tesla Optimus has published no model details. These four are from prior knowledge and were not re-verified today.
E. Benchmark results across models
LIBERO (Spatial / Object / Goal / Long; average) — saturated
Almost all numbers are self-reported. "XR0 table" means compiled in the Xiaomi-Robotics-0 paper. The standard protocol is 50 trials per task (prior knowledge); LeRobot re-runs use only 10 per task.
| Model | Spatial | Object | Goal | Long | Average | Source |
|---|---|---|---|---|---|---|
| DM0.5 | — | — | — | — | 99.0 | Dexmal README |
| Galaxea G0.5 | — | — | — | — | 98.9 | README |
| Xiaomi-Robotics-0 | 98.8 | 100 | 98.8 | 97.2 | 98.7 | paper |
| VLA-Adapter-Pro | 99.6 | 99.6 | 98.2 | 96.4 | 98.5 | README |
| AgiBot GO-2 | — | — | — | — | 98.5 | press; closed model |
| EO-1 | 99.7 | 99.8 | 99.2 | 94.8 | 98.2 | XR0 table |
| X-VLA | — | — | — | — | 98.1 | README |
| RIPT-VLA | — | — | — | — | 97.5 | LIBERO-Plus paper |
| π0.5 (LeRobot re-run) | 97.0 | 99.0 | 98.0 | 96.0 | 97.5 | LeRobot docs |
| OpenVLA-OFT | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 | paper |
| π0.5 (openpi) | 98.8 | 98.2 | 98.0 | 92.4 | 96.85 | LeRobot docs |
| FLOWER | 97.5 | 99.1 | 96.1 | 94.9 | 96.9 | XR0 table |
| MemoryVLA | 98.4 | 98.4 | 96.4 | 93.4 | 96.7 | XR0 table |
| GR00T N1.7 (LeRobot) | 95 | 100 | — | — | 96.5 | HF/NVIDIA blog |
| VLA-JEPA (LeRobot) | 95 | 100 | 98 | 93 | 96.5 | LeRobot docs |
| Discrete Diffusion VLA | 97.2 | 98.6 | 97.4 | 92.0 | 96.3 | XR0 table |
| InternVLA-M1 | 98.0 | 99.0 | 93.8 | 92.6 | 95.9 | paper |
| UniVLA (OpenDriveLab) | 96.5 | 96.8 | 95.6 | 92.0 | ≈95.2 | README |
| π0.5-KI | 98.0 | 97.8 | 95.6 | 85.8 | 94.3 | InternVLA-M1 table |
| π0 | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 | XR0 table |
| GR00T-N1 | 94.4 | 97.6 | 93.0 | 90.6 | 93.9 | XR0 table |
| NORA | — | — | — | — | 87.9 | LIBERO-Plus paper |
| GR00T N1.5 (LeRobot) | — | — | — | — | 87 | HF/NVIDIA blog |
| SmolVLA | — | — | — | — | ~87.3 | paper; not re-verified |
| π0-FAST | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 | XR0 table |
| Dita | — | — | — | — | 82.4 | README |
| WorldVLA | — | — | — | — | 79.1 | LIBERO-Plus paper |
| SpatialVLA | — | — | — | — | 78.1 | README |
| OpenVLA | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | XR0 table |
| Octo | — | — | — | — | 75.1 | X-VLA table |
Takeaways:
- At least seven models are at or above 98%, so LIBERO no longer separates the leaders.
- The LeRobot re-run of π0.5 came out 0.65 points higher than PI's own number, with only 10 trials per task, which shows the protocol noise.
Robustness: LIBERO-Plus and LIBERO-PRO (independent)
LIBERO-Plus (NUS, Fudan, Tongji, Shanghai Innovation Institute; CVPR 2026). Success rates after perturbation:
| Model | Original | Camera | Robot start | Language | Light | Background | Noise | Layout |
|---|---|---|---|---|---|---|---|---|
| OpenVLA | 76.5 | 1.1 | 4.1 | 26.8 | 4.4 | 25.3 | 19.3 | 31.6 |
| OpenVLA-OFT | 97.1 | 59.7 | 37.2 | 81.5 | 85.8 | 92.4 | 76.7 | 77.1 |
| π0 | 94.2 | 15.8 | 6.6 | 61.0 | 79.6 | 78.5 | 79.4 | 70.4 |
| π0-FAST | 85.5 | 66.4 | 24.8 | 63.3 | 73.0 | 67.7 | 75.8 | 70.3 |
| NORA | 87.9 | 4.0 | 41.1 | 67.0 | 31.0 | 50.5 | 17.6 | 63.9 |
| UniVLA | 95.2 | 4.3 | 50.3 | 71.8 | 59.1 | 80.0 | 25.3 | 34.3 |
| WorldVLA | 79.1 | 0.3 | 30.2 | 44.2 | 29.4 | 14.5 | 12.2 | 39.4 |
| RIPT-VLA | 97.5 | 58.3 | 36.7 | 80.1 | 87.9 | 90.4 | 73.8 | 76.5 |
Findings from the paper:
- Models are "largely insensitive to language variations… tend to ignore language instructions completely".
- Camera viewpoint changes cause collapse.
- Models memorize object positions.
- Wrist cameras help.
GO-2 claims 86.6 on LIBERO-Plus zero-shot (self-reported).
LIBERO-PRO (Harvard, MIT et al.) tested OpenVLA, π0 and π0.5. Under changes to objects' starting positions, success drops to 0%. Rephrased instructions and task-level changes cause near-total failure, and the models replay "nearly identical trajectories", which indicates memorization.
SimplerEnv (Google Robot visual matching and variant aggregation; WidowX)
All self-reported, compiled from the sources named.
| Model | Google Robot VM | Google Robot VA | WidowX |
|---|---|---|---|
| Xiaomi-Robotics-0 | 85.5 | 74.7 | 79.2 |
| X-VLA | 83.5 | 76.4 | 95.8 |
| InternVLA-M1 | 80.7 | 76.0 | 71.7 |
| SpatialVLA (InternVLA-M1 table) | 75.1 | 70.7 | — |
| CogACT | 74.8 | 61.3 | 51.3 |
| RT-1 | 52.4 | — | — |
| RT-2-X | 46.3 | 54.4 | — |
| Galaxea G0.5 | — | — | 87.3 (Bridge) |
| EO-1 | — | — | 72.7 |
| OpenVLA-OFT (X-VLA table) | — | — | 63.0 |
| GR00T N1.5 | — | — | 61.9 |
| π0 (X-VLA table) | — | — | 58.8 |
| π0-FAST | — | — | 48.3 |
| UniVLA | — | — | 47.9 |
| GR00T-N1 | — | — | 45.0 |
| Octo | — | — | 16.8 |
SpatialVLA also reports, zero-shot, 71.9 on Google Robot and 34.4 on WidowX (README). PI does not report SimplerEnv, and the π0 numbers come from third-party reproductions.
CALVIN (average number of chained tasks completed, maximum 5)
- ABC→D: Xiaomi-Robotics-0 4.75; FLOWER 4.54; VLA-Adapter-Pro 4.50; X-VLA 4.43.
- ABCD→D: Xiaomi 4.80; FLOWER 4.67; "UniVLA" 4.63 (the Xiaomi table may mean BAAI's "UniVLA", not OpenDriveLab's; ambiguous); MDT 4.52; RoboVLMs 4.49; MoDE 4.39; GR-1 4.21; RoboFlamingo 4.09.
- All self-reported. The benchmark is close to saturation.
RoboTwin 2.0 (50 dual-arm tasks; 731 objects; 5 embodiments; AgileX ALOHA arms)
- Leaderboard rule: submissions need public code, public weights and a technical report.
- Self-reported results: DM0.5 93.6 clean / 93.3 randomized; Galaxea G0.5 93.3; InternVLA-A1 89.40 easy / 89.64 hard; LingBot-VLA (with depth) 88.56 / 86.68; X-VLA 70.
- Synthetic-data claim from the RoboTwin 2.0 paper: synthetic data plus 10 real demonstrations improved success by 367% relative; synthetic-only zero-shot by 228%.
RoboChallenge Table30 (real robots: UR5, Franka, Cobot Magic ALOHA, ARX-5; run by Dexmal and Hugging Face)
-
Official baselines (Oct 2025 technical report):
Model Success rate Progress score π0.5 (task-specific) 43.7% 62.2 π0 (task-specific) 28.3% 47.6 CogACT 11.7% 21.8 π0.5 multi-task generalist 17.7% 31.3 π0 multi-task generalist 9.3% 20.6 -
Aggregator snapshot (sota2, about Apr 2026): DM0 62%; GigaBrain 51.67%; Spirit-v1.5 51%; π0.5 42.67%; WALL-OSS 35.33%; "A1" 29%; π0 28.33%; X-VLA 21.33%; RDT-1B 15%.
-
Table30v2: DM0.5 scores 54.42 with 43.0% success (self-reported).
-
Caveat: Dexmal co-runs the benchmark and builds the leading DM0 model.
RoboArena (distributed real-robot evaluation on DROID setups; Berkeley, Stanford, UW, NVIDIA, Penn, UT Austin and others)
- Protocol: double-blind pairwise A/B trials ranked with a Bradley-Terry model.
- Scale in paper v2 (Nov 29, 2025): 4,284 episodes, 612 pairwise comparisons, 7 institutions, 7 policies. π0-FAST-DROID ranked highest, and "discrete action tokenization outperform[s] diffusion" in language-conditioned evaluations.
- Later "#1 on RoboArena" claims (both self-reported; I could not read the live leaderboard): GR00T N2 (Mar 2026) and Cosmos3-Nano-Policy (May 2026).
BEHAVIOR Challenge (Stanford)
- 2025 (50 household tasks): the winning team (Larchenko, Zarin, Karnatak) built on π0.5 and scored 26% q-score on both public and private leaderboards; the solution is open-sourced. "Openpi Comet" (arXiv 2512.10071) is another top write-up. Galaxea G0.5 reports 0.3136 on BEHAVIOR-1K (self-reported).
- 2026: launched Jul 2, 2026; submissions due Oct 16; winners Nov 4. 100 full-length tasks, with 20,000 teleoperated demonstrations totalling 1,950 h provided. Baselines are π0.5 and GR00T N1.7.
Other 2026 evaluations and humanoid-specific gaps
- New benchmarks named in 2026 sources: RoboChallenge Table30v2, RoboDojo-Sim, RoboColiseum, VLA-Arena, MolmoSpaces (named by NVIDIA), GM-100 (LingBot), RoboLab (Cosmos 3 card) and Genie Sim 3.0 (AgiBot). Results across them are mostly self-reported, and the field is fragmenting.
- Humanoid evaluation: there is no standardized open humanoid manipulation benchmark.
- NVIDIA reports GR00T on 24 simulated GR-1 "Digital Cousin" tasks, 24 RoboCasa tasks and 9 DexMG tasks; numbers not captured.
- Gemini Robotics 2 reports whole-body pickup on Apollo 2 at 45.7–76.3%.
1. Also notable (brief)
- GR00T N2 (NVIDIA): world-action model previewed Mar 16, 2026; not yet released. nvidianews link above.
- Cosmos-Reason2 / Transfer2.5 / cosmos-rl (NVIDIA): ACTIVE. github.com/nvidia-cosmos.
- Giga-World-Policy-0.5 (GigaAI): ACTIVE (Jul 30, 2026). HF open-gigaai/Giga-World-Policy-0.5.
- FastWAM, LaWAM, LingBot-VA (world-action models) and EVO1 (VLA): ACTIVE, recently added to LeRobot; details not verified. huggingface.co/docs/lerobot.
- Multitask DiT Policy and the reward models SARM, Robometer, TOPReward: ACTIVE, in LeRobot.
- LingBot-World-V2 / LingBot-Video (Robbyant): ACTIVE (Jul 2026). HF robbyant.
- Dexbotic (Dexmal VLA toolbox; 1,389 stars; MIT; updated Aug 6, 2026), realtime-vla (604 stars), opendw world model (219 stars): ACTIVE. github.com/Dexmal.
- WALL-WM (X Square Robot): ACTIVE (May 2026). github.com/X-Square-Robot/wall-x.
- G0Tiny (Galaxea): 250M edge model, Feb 2026.
- Being-H0.5 / H0.7: newer iterations of Being-H0; details unverified.
- NORA-1.5: released; date unverified.
- EO-1 continuation at SHAILAB-IPEC/EO1: not fetched.
- GraspMolmo (AI2): HF allenai/GraspMolmo, May 2026.
- PhysBrain 1.5 and RynnBrain (Alibaba): GGUF quantizations seen on HF in 2026; origin not verified.
- SimVLA: HF YuankaiLuo/SimVLA-LIBERO, Feb 2026; not verified.
- MemoryVLA (96.7) and Discrete Diffusion VLA (96.3): LIBERO results cited in other papers' tables.
- RIPT-VLA: LIBERO 97.5, robust on LIBERO-Plus.
- WorldVLA: superseded by RynnVLA-002.
- NVIDIA Alpamayo-R1 / 1.5: driving VLAs that inflate HF's "robotics" download counts (55.5k a month); out of scope.
- Older baselines, DORMANT: MiniVLA, TinyVLA, RoboVLMs, Seer, VPP, MDT, MoDE, GR-1, RoboFlamingo.
2. Comparison table
HF downloads are per month from the model card unless marked (L) for a listing figure. "NC" means non-commercial.
| Model | Org | Params | Released | License: code / weights (commercial?) | Action head | LIBERO avg | SimplerEnv (VM / VA; WidowX) | Other key result | HF downloads / month |
|---|---|---|---|---|---|---|---|---|---|
| π0 | Physical Intelligence | ~3.3B | 2024-10 (open 2025-02) | Apache-2.0 / Gemma terms (yes, with conditions) | flow matching | 94.2 | –; 58.8 (3rd-party) | RoboChallenge 28.3% | 12–19k (L) |
| π0-FAST | PI | ~3B | 2025-01/02 | same | autoregressive FAST tokens | 85.5 | –; 48.3 | Top of RoboArena paper (DROID) | 1.9k (L) |
| π0.5 | PI | ~3B+ | 2025-04 (open 2025-09) | same | FAST-token pretraining, flow expert | 96.85 / 97.5 (LeRobot) | – | RoboChallenge 43.7%; BEHAVIOR '25 winning base | 19.4k |
| GR00T N1.5 | NVIDIA | 3B | 2025 | Apache / NVIDIA (no) | flow DiT | 87 (LeRobot) | –; 61.9 | – | 1.4k |
| GR00T N1.6 | NVIDIA | 3B | ~2025-12 | Apache / OneWay NC (no) | flow DiT | – | – | – | 29.2k |
| GR00T N1.7 | NVIDIA | 3B | 2026-04 early access | Apache / Open Model License (yes) | flow DiT | 96.5 (LeRobot) | – | BEHAVIOR '26 baseline | 147k (card) / 45k (L) |
| OpenVLA | Stanford et al. | 7.5B | 2024-06 | MIT / MIT + Llama 2 (yes, with conditions) | autoregressive bins | 76.5 | weak | LIBERO-Plus camera 1.1 | 460k (card) / 169k (L) |
| OpenVLA-OFT | Stanford | 7.5B | 2025-02 | MIT | parallel decoding, L1 | 97.1 | –; 63.0 | Most robust on LIBERO-Plus | 13.9k (L, one checkpoint) |
| Octo-Base | Berkeley et al. | 93M | 2024-05 | MIT (yes) | diffusion | 75.1 | –; 16.8 | – | 115 |
| SmolVLA | Hugging Face | 450M | 2025-06 | Apache (yes) | flow matching | ~87.3 (unverified) | – | SO-100 78.3%; 7,808 fine-tunes | 159k |
| RDT-1B | Tsinghua | 1.2B | 2024-10 | MIT (yes) | diffusion | – | – | RoboChallenge 15% | 405 |
| RDT2-VQ / FM | Tsinghua | ~8B | 2025-09 | Apache (yes) | RVQ autoregressive / flow | – | – | Zero-shot on unseen embodiments | 371 (card) / 13.7k (L) |
| X-VLA | Tsinghua AIR / Shanghai AI Lab | 0.9B | 2025-10 | Apache (yes) | flow matching | 98.1 | 83.5 / 76.4; 95.8 | CALVIN 4.43; RoboChallenge 21% | 10.5k |
| CogACT | MSRA | ~7B | 2024-11 | MIT (yes) | DiT | – | 74.8 / 61.3; 51.3 | RoboChallenge 11.7% | n/a |
| SpatialVLA | Shanghai AI Lab et al. | 4B | 2025-01 | MIT + Gemma | autoregressive spatial tokens | 78.1 | 71.9 (zero-shot); 34.4 | – | n/a |
| UniVLA | OpenDriveLab | 7B | 2025-05 | Apache (yes) | latent action + decoder | ~95.2 | –; 47.9 | – | n/a |
| WALL-OSS | X Square | 4B | 2025-09; 0.5 in 2026-05 | Apache (yes) | flow / FAST | – | – | RoboChallenge 35% | 908 |
| GO-1 | AgiBot | 3B | open 2025-09 | CC BY-NC-SA (no) | latent planner + diffusion | – | – | – | 149 |
| InternVLA-M1 | Shanghai AI Lab | 4.1B | 2025-10 | MIT (yes) | DiT (86M) | 95.9 | 80.7 / 76.0; 71.7 | – | n/a |
| InternVLA-A1 | Shanghai AI Lab | 3B | 2026-01 | ? / CC BY-NC-SA (no) | mixture of transformers with generation expert | – | – | RoboTwin2 89.4 / 89.6 | 81 |
| MolmoAct2 | AI2 | 5B | 2026-05 | Apache (yes) | flow expert + OpenFAST | n/c | n/c | Claims > π0.5 | 19.8k |
| Xiaomi-Robotics-0 | Xiaomi | 4.7B | 2026-02 | Apache (yes) | flow DiT | 98.7 | 85.5 / 74.7; 79.2 | CALVIN 4.75 | 3.3k (L) |
| Spirit v1.5 | Spirit AI | 5B | 2026-01 | MIT / Apache (yes) | DiT | – | – | RoboChallenge #1 at launch (~51%) | 131 |
| GigaBrain-0.7 | GigaAI | ~3.5B | 2026-08 | Apache (yes) | three-system | – | – | GigaBrain 51.7% RoboChallenge | 6.0k |
| DM0.5 / OpenDM | Dexmal | n/d | 2026-07 | Apache (yes) | n/d | 99.0 | – | RoboTwin2 93.6; Table30v2 43% | n/a |
| LingBot-VLA | Ant / Robbyant | 4B | 2026-01 | Apache (yes) | n/d | – | – | RoboTwin2 88.6; GM-100 17% | n/a |
| Galaxea G0.5 | Galaxea | ~2B | 2026-06 | NC (no) | autoregressive tokens | 98.9 | –; 87.3 | RoboTwin2 93.3; BEHAVIOR 0.31 | n/a |
| EO-1 | Shanghai AI Lab | 3B | 2025-08 | MIT (yes) | autoregressive + flow | 98.2 | –; 72.7 | – | 55 |
| VLA-Adapter-Pro | OpenHelix | ~0.5B | 2025-09 | MIT (yes) | adapter policy | 98.5 | – | CALVIN 4.50 | n/a |
| FLOWER | KIT | ~1B | 2025 | MIT (yes) | rectified flow | 96.9 | – | CALVIN 4.54 | n/a |
| VLA-JEPA | academic | ~2B+ | 2026-02 | Apache (yes) | flow DiT + JEPA loss | 96.5 | – | – | 2.8k (L) |
| Cosmos3-Nano-Policy | NVIDIA | 16B | 2026-05-31 | OpenMDW-1.1 (yes) | autoregressive + diffusion | – | – | Claims RoboArena #1 | 2.5k |
| DreamZero | NVIDIA GEAR | 14B | 2026-02/03 | CC-BY-NC (no) | video-diffusion world-action model | – | – | ~7 Hz; >2× generalization (self) | 1.4k |
3. Gaps and pain points, with evidence, and what a data / sim / environments company could supply
- Commercially licensable pretraining data is the binding constraint.
- NVIDIA staff on GR00T: weights were research-only "due to pre-training data license constraints… a data issue, not a license choice" (Jul 16, 2025).
- Galaxea moved G0.5 to a non-commercial license in 2026.
- AgiBot World data and GO-1 weights are CC BY-NC-SA, as are InternVLA-A1's weights; DreamZero is CC-BY-NC.
- → Opportunity: consented, commercially clean, multi-embodiment datasets with clear provenance.
- Human and egocentric data is now a first-class pretraining source, but it needs action labels.
- GR00T N1.7 uses 20K h of EgoScale human video in a relative end-effector action space shared with robots.
- RDT2 uses 10k+ h of UMI data. OpenWAM's mix is 30% egocentric human. InternVLA-A1 and UniVLA use human video.
- Being-H0 models hands with MANO.
- → Opportunity: egocentric capture with calibrated 6D wrist poses, hand and finger pose (gloves), and synchronized language and subtask annotations, delivered in LeRobot v3.0 format (which GR00T-in-LeRobot and MolmoAct2 both use).
- Dexterous hands are under-served.
- Every mainstream benchmark (LIBERO, SimplerEnv, CALVIN, RoboTwin, RoboChallenge) uses parallel grippers.
- Only GR00T N1.7 (finger-level control) and Being-H0 target hands in the open. Gemini Robotics 2 reaches only 32–92% on five-finger tasks.
- Users ask whether GR00T can learn suction-gripper actions (issue, Jul 15, 2026).
- → Opportunity: dexterous teleoperation and glove datasets, and a hand-manipulation benchmark.
- Humanoid whole-body data and evaluation are scarce.
- Open humanoid VLAs: GR00T on G1 and GR-1, and UnifoLM-VLA-0 on G1.
- Benchmarks: no standard; Gemini's whole-body pickup is 45.7–76.3%.
- → Opportunity: whole-body teleoperation corpora and humanoid simulation task suites.
- Robustness data matters more than more of the same demonstrations.
- LIBERO-Plus: under camera changes π0 falls to 15.8% and OpenVLA to 1.1%; under robot-start changes π0 falls to 6.6%; models ignore language.
- LIBERO-PRO: object position changes drop success to 0%.
- RoboTwin 2.0's domain randomization gave large relative gains.
- → Opportunity: procedurally varied simulated environments (cameras, starting states, layouts, paraphrased instructions) and "stress-test" evaluation suites.
- Evaluation is expensive, fragmented and often conflicted.
- Cost: RoboArena needed 4,284 episodes across 7 institutions; TRI used 1,800 real rollouts plus 47k simulated ones.
- Conflict: Dexmal both runs RoboChallenge and tops it.
- Saturation: LIBERO is saturated, and LeRobot uses only 10 trials per task.
- Disputed claims: GR00T N2 and Cosmos 3 each claim RoboArena #1.
- → Opportunity: neutral evaluation-as-a-service, both real-robot fleets and world-model-based evaluation.
- Long-horizon tasks and generalists are weak.
- BEHAVIOR 2025 winner: 26%.
- RoboChallenge: π0.5 drops from 43.7% task-specific to 17.7% as a generalist.
- LingBot scores 17% on GM-100.
- BEHAVIOR 2026 supplies 1,950 h of demonstrations for 100 tasks.
- → Opportunity: long-horizon demonstrations with subtask segmentation, and reward or progress labels. The reward models SARM, Robometer and TOPReward, and RoboBrain 2.5's value estimation, all consume this kind of label.
- Small, high-quality fine-tuning sets are what customers actually buy.
- SmolVLA needs at least 50 episodes ("25… not enough").
- TRI: pretraining cuts the data needed 3–5×.
- Gemini On-Device 2 adapts with fewer than 200 examples.
- → Opportunity: per-embodiment, per-task "fine-tune packs" plus collection tooling.
- Inference cost and latency.
- GR00T N1.7 runs at 4.6 Hz on Orin versus about 36 Hz on H100.
- DreamZero (14B) runs at about 7 Hz; Cosmos3 policies are 16B; CogACT takes 181 ms.
- Users struggle with TensorRT export (openpi #982) and with fine-tuning on 24 GB GPUs (GR00T issue).
- Real-time action chunking (PI's real-time chunking work, LeRobot's RTC mode, Dexmal's realtime-vla) is the workaround.
- Tooling friction.
- Data conversion: "convert_aloha_data_to_lerobot.py is outdated and extremely memory-heavy" (openpi #986).
- Backend gaps between PyTorch and JAX (#989).
- Checkpoint and config confusion in GR00T issues.
- → Opportunity: data pipelines and format converters as a product.
4. Every GitHub repo and HF model ID referenced, with counts fetched 2026-10-02
GitHub (stars / forks):
- Physical-Intelligence/openpi 13.2k / 2.3k (239 open issues)
- NVIDIA/Isaac-GR00T 8.2k / 1.5k (213 open issues)
- huggingface/lerobot 27.8k / 5.7k
- openvla/openvla 6.8k / 820
- moojink/openvla-oft 1.3k / 195
- octo-models/octo 1.7k / 281
- thu-ml/RoboticsDiffusionTransformer 1.8k / 163
- thu-ml/RDT2 805 / 58
- 2toinf/X-VLA 728 / 71
- X-Square-Robot/wall-x 1.3k / 99
- OpenDriveLab/AgiBot-World 3.2k / 218
- AgibotTech/Genie-Envisioner 586 / 27
- InternRobotics/InternVLA-M1 432 / 27
- InternRobotics InternVLA-A-series (exact repo name not captured) 479 / 34
- allenai/molmoact 383 / 43
- allenai/molmoact2 782 / 63
- FlagOpen/RoboBrain2.0 1.1k / 116
- BeingBeyond/Being-H0 59 / 0
- EO-Robotics/EO1 291 / 29 (archived)
- microsoft/Magma 1.9k / 164
- alibaba-damo-academy/RynnVLA-002 1.1k / 67
- OpenGalaxea/GalaxeaVLA 759 / 57
- unitreerobotics/unifolm-vla 546 / 57
- unitreerobotics/unifolm-world-model-action 1.1k / 140
- nvidia-cosmos/cosmos-predict2.5 1,373; cosmos-transfer2.5 743; cosmos-reason2 452; cosmos-reason1 962; cosmos-predict1 474; cosmos-transfer1 827; cosmos-cookbook 477; cosmos-rl 474
- facebookresearch/vjepa2 4.7k / 580
- real-stanford/diffusion_policy 4.6k / 848
- tonyzhaozh/act 2.2k / 420
- YanjieZe/3D-Diffusion-Policy 1.4k / 171
- liruiw/HPT 542 / 36
- OpenHelix-Team/VLA-Adapter 2.3k / 208
- declare-lab/nora 222 / 21
- intuitive-robots/flower_vla_calvin 94 / 18
- RoboDita/Dita 172 / 9
- microsoft/CogACT 432 / 40
- SpatialVLA/SpatialVLA 727 / 51
- OpenDriveLab/UniVLA 1.1k / 69
- Spirit-AI-Team/spirit-v1.5 624 / 36
- XiaomiRobotics/Xiaomi-Robotics-0 641 / 70
- Dexmal/opendm 2.2k / 184 (the org page showed 304)
- Dexmal/dexbotic 1,389; realtime-vla 604; realtime-vla-v2 145; realtime-vla-flash 102; opendw 219; dexbotic-benchmark 39
- Robbyant/lingbot-vla (stars not captured)
Hugging Face (downloads last month on the card unless (L); likes):
- openvla/openvla-7b 460,364 (169k L); 260
- lerobot/smolvla_base 159,287 (55–71k L); 445; 7,808 fine-tunes
- lerobot/pi05_base 19,354; 103; 633 fine-tunes
- lerobot/pi0_base 12.3–18.8k (L); 26
- lerobot/pi0fast-base 1.86k (L); 26
- lerobot/pi05_libero_finetuned_v044 14–17k (L)
- lerobot/pi05_libero_base 2.9k (L)
- lerobot/smolvla_libero 9.9–19.5k (L)
- HuggingFaceVLA/smolvla_libero 15.7–16.9k (L)
- lerobot/xvla-libero 2.9–4.1k (L)
- lerobot/VLA-JEPA-Pretrain 2.78k (L)
- nvidia/GR00T-N1.7-3B 147,412 (45k L); 141
- nvidia/GR00T-N1.6-3B 29,218; 91
- nvidia/GR00T-N1.5-3B 1,402; 197
- nvidia/GR00T-N1.7-SimplerEnv-Bridge 532
- nvidia/GR00T-H-N1.7 (seen in search results; could not open)
- GEAR-Dreams/DreamZero-DROID 1,423; 34
- nvidia/Cosmos3-Nano-Policy-DROID 2,472; 34
- nvidia/Cosmos3-Edge 950k (L); Cosmos3-Nano 282k (L); Cosmos3-Super 121k (L); Cosmos3-Super-Image2Video 74.8k (L)
- nvidia/Cosmos-Predict2.5-2B 7,631; 167
- allenai/MolmoAct2 19,785; 24
- allenai/MolmoAct2-LIBERO 4.1–20.9k (L); MolmoAct2-SO100_101 6.76k (L); -Think 3.94k (L); -Pretrain 3.42k (L); -BimanualYAM 3.15k (L); -DROID 1.72k (L)
- moojink/openvla-7b-oft-finetuned-libero-spatial 13.9k (L); -object 6.38k (L); -10 4.33k (L); -goal 3.46k (L); -spatial-object-goal-10 2.14k (L)
- robotics-diffusion-transformer/RDT2-VQ 371 (13.7k L); 22
- robotics-diffusion-transformer/rdt-1b 405; 106
- 2toINF/X-VLA-Pt 10,471; 14
- 2toINF/X-VLA-Libero 2.45k (L)
- rail-berkeley/octo-base-1.5 115; 18
- agibot-world/GO-1 149; 20
- IPEC-COMMUNITY/EO-1-3B 55; 14
- InternRobotics/InternVLA-A1-3B 81; 40
- Spirit-AI-robotics/Spirit-v1.5 131; 31
- open-gigaai/GigaBrain-0.7-3.5B-Base 6,005; 12
- open-gigaai/Giga-World-Policy-0.5 3.58k (L)
- x-square-robot/wall-oss-flow 908; 34
- XiaomiRobotics/Xiaomi-Robotics-0-LIBERO 3.34k (L); 14
- declare-lab/nora-long 3.83k (L)
- OpenWAM/OpenWAM-Alpha-Pretrain-Foundation-Model 343
- robbyant/lingbot-vla-v2-6b-robotwin (7 likes)
- lerobot/diffusion_pusht 3.5–4.0k (L); 64
- lerobot/act_aloha_sim_transfer_cube_human 3.0k (L)
- Referenced but not fetched: unitreerobotics/UnifoLM-VLA-Base; BAAI/RoboBrain2.5-8B-NV and -4B; BAAI/RoboBrain2.0-7B and -32B; CogACT/CogACT-Base; IPEC-COMMUNITY/spatialvla-4b-224-pt; OpenGalaxea/G05; agibot-world/GO-1-Air.
- HF tag totals: pipeline_tag=robotics has 26,403 models; the lerobot tag has 23,950 models.
Benchmark and other sources opened:
- LIBERO-Plus: https://arxiv.org/html/2510.13626v1
- LIBERO-PRO: https://arxiv.org/html/2510.03827v1
- RoboChallenge report: https://arxiv.org/html/2510.17950v1
- RoboChallenge aggregator snapshot: https://www.sota2.com/research/sota/robotic-manipulation-on-robochallenge-table30
- RoboArena paper: https://arxiv.org/html/2506.18123v2
- BEHAVIOR 2025 winner: https://huggingface.co/papers/2512.06951
- BEHAVIOR 2026: https://behavior.stanford.edu/challenge
- RoboTwin 2.0 paper: https://arxiv.org/abs/2506.18088
- RoboTwin leaderboard (protocol only): https://robotwin-platform.github.io/leaderboard
- Xiaomi-Robotics-0 tables: https://arxiv.org/html/2602.12684v1
- InternVLA-M1 tables: https://arxiv.org/html/2510.13778
- X-VLA tables: https://arxiv.org/html/2510.10274
- Gemini Robotics 2: https://www.marktechpost.com/2026/07/30/google-deepmind-gemini-robotics-2-whole-body-control-dexterity-multi-robot-collaboration/
- Generalist AI blog: https://generalistai.com/blog
- AgiBot GO-2: https://www.therobotreport.com/agibot-releases-go-2-foundation-model-embodied-ai/
- MolmoAct2 coverage: https://siliconangle.com/2026/05/05/ai2-releases-molmoact-2-enhancing-robot-intelligence-real-world/