RoboMIND

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

How to read this picture

RoboMIND is a dataset of about 107,000 demonstrations on four robot types, used to train robot control models (policies). About 28% of the demonstrations come from simulation, and its paper tests policies on real robots.12

Sources
Last checked 10 Oct 2026Full entry40 of 53 facts checked at the sourceNext check 8 Apr 2027
Runs in
Real robots1
Checked against real robots
Real robots
Skill
Handling objects
Robot
One arm, Two arms, Humanoid, Robot hand1
Used by
Used as pretraining data for EO-1 and XR-1341+1
244 citations
Licence
Allowed56
Commercial use: allowed

What a score here does not tell you Inferred

  1. Whether a policy would get the same score in another lab.All tests in the paper ran on X-Humanoid's own robot setups.
  2. How well a policy does in new surroundings.In the one task tested this way, new tablecloths cut success to 0 to 2 of 10 trials.
  3. Whether the digital twin (the simulated copy) predicts real-robot results.The authors checked it with only 2 policies on 5 tasks.

Comparisons with real robots

StudyResultWhat was comparedDone by
RoboMIND digital-twin check (Table VII)
May 2025
Pearson correlation r = 0.83 (ACT) and 0.91 (Diffusion Policy), each across 5 tasks17ACT and Diffusion Policy, trained on real data, were each run 10 times on the same 5 Franka tasks in the Isaac Sim twin and on the real robot. The study’s authors described the result as “positive correlations”.The benchmark’s authors
RoboMIND co-training study (Figure 17)
Feb 2025
No statistic was reported. The policy trained only on twin data scored 0.9 in the twin and 0.1 on the real robot. The policy trained only on real data scored 0 in the twin and 0.6 on the real robot.877 ACT policies trained on different mixes of real and twin data were each scored on one task (FR-UprightBlueCup) in the twin and on the real robot. The study’s authors described the result as “simulation data alone is insufficient”.The benchmark’s authors

Our assessment Opinion

RoboMIND is mainly training data. Its test results are one-off tests by its builders.

Reasoning

RoboMIND is mainly training data. Its paper's success rates are one-off tests by the builders on their own robots and task setups, with 10 trials each and no error bars. They are not comparable with results in other papers.

Confidence: high

The check of the digital twin is too small to show that the twin can be used as a test.

Reasoning

The digital-twin check is small. Two policies on five tasks, with ten trials each, cannot show that the twin ranks many policies correctly. The twin also scored policies far lower than the real robot. A policy trained only in the twin scored well there and failed on the real robot.

Confidence: high

When quoting its size, name the version and say whether simulated data is counted.

Reasoning

When quoting its size, name the version and say whether simulation is included. Version 1.0 has 55k trajectories, and versions 1.1 and 1.2 have 107k, of which about 30k are simulated. RoboMIND 2.0, a different dataset, has over 310K.

Confidence: high

Known problems 6

  1. The name is shared with RoboMIND 2.0

    The name 'RoboMIND' is often used for RoboMIND 2.0, a separate and larger dataset. Their sizes and download counts are easy to mix up.2910

    Details

    The version 1 card and site announce 'RoboMIND V2.0', a different and larger dataset (over 310K trajectories, six robot types). Chinese news coverage uses 'RoboMIND' for 2.0 figures (300,000+ trajectories, 700+ tasks, over 20 million downloads). Sizes and download counts are easy to mix up.

  2. About 28% of the 'real-world' data is simulated

    The site and dataset card describe '107k real-world' trajectories. By the paper's own breakdown, about 30,000 of them come from simulation.1121+1

    Details

    The project site and dataset card describe '107k real-world demonstration trajectories'. The paper's own breakdown counts 30,035 of them from the Isaac Sim digital twin, and Franka's 49.2% share includes over 26,070 twin trajectories. The paper's comparison table places RoboMIND among real-world datasets. Version 1.0 also mixed in 11,783 simulated trajectories while its abstract said '55k real-world'.

  3. Trajectory counts differ between versions, the paper's figure and its text

    The number of trajectories differs between versions, between the paper's text and its figure, and on the dataset card. Some of these differences are still unexplained.1217+1

    Details

    Version 1.0 (arXiv v1) has 55k trajectories and 308.6 hours; versions 1.1 and 1.2 have 107k trajectories but 305.5 hours. Within the RSS paper, the text lists Franka 26,856, Tien Kung 15,187, AgileX 10,269, UR5e 25,170 and simulation 30,035, while Figure 1(a) and the card list Franka 52,926, humanoid 19,152, AgileX 10,629 and UR5e 25,170. Most of the gap is explained by Figure 1 and the card folding twin data into each robot (see facts.demonstrations). Still unresolved: AgileX 10,629 vs 10,269; Franka real 26,856 (introduction) vs 26,866 (Section IV-A); object classes 61 vs 69 for v1.0; and the drop in hours from 308.6 to 305.5 while trajectories doubled.

  4. End-effector data is frozen in some subsets

    X-Humanoid's own checks found that end-effector values (the position and angle of the robot's gripper or hand) stay almost constant in three subsets. Users report further problems.6213+1

    Details

    X-Humanoid's data-check repository (2026-04) reports end-effector values that stay almost constant in three subsets: h5_ur_1rgb, h5_simulation and h5_sim_franka_3rgb, probably from a blocking problem during collection. It recommends using joint positions and recomputing end-effector poses with forward kinematics. Users report constant UR end-effector poses (discussion #8) and UR episodes whose joints barely move although the video shows motion (discussion #6); one user posted a scripted count of 22,987 of 25,721 UR episodes (89.37%) with the first six joint values unchanged, which X-Humanoid has not answered. The card also notes BGR image order in four subsets and RGB in the rest, and 675 Franka trajectories with only two of three cameras.

  5. The digital twin check is small and cannot be repeated

    The digital twin was checked with 2 policies on 5 tasks. The twin is not released, so others cannot repeat the check.17

    Details

    Table VII compares 2 policies on 5 Franka tasks with 10 trials per task; each Pearson r rests on 5 points and has no interval. The twin scored policies far lower than the real robot (ACT 1.4/10 vs 5.4/10 on average). In Figure 17 a policy trained only on twin data scored 0.9 in the twin and 0.1 on the real robot, which the authors attribute to physics gaps in contact-rich tasks. The version 1 twin is not released, so others cannot rerun the check.

  6. The backgrounds are simple and fixed

    In the authors' test, three new tablecloths cut success on one task to 0 to 2 of 10 trials.1

    Details

    The authors list relatively simple background environments as a limitation. In their generalisation test, three unseen tablecloths cut FR-PlaceBreadPlate success to 0-2 of 10 for OpenVLA, RDT-1B and CrossFormer (from 4, 9 and 10 of 10).

Details

About

What it is
Dataset Inferred12
More

The title says 'Benchmark', and the paper says RoboMIND 'serves as a benchmark' for the methods it tests. What is released is training data plus conversion and training code; the real-robot test setups are not released.

Built by
X-Humanoid, with Peking University and BAAI11511
More

37 authors on the RSS page.

Beijing Innovation Center of Humanoid Robotics (X-Humanoid) · 北京人形机器人创新中心有限公司. Builds the Tien Kung humanoid. Corresponding author Jian Tang; project leaders Zhengping Che and Xiaozhu Ju.116

Peking University · State Key Laboratory of Multimedia Information Processing, School of Computer Science. Corresponding author Shanghang Zhang.1

Beijing Academy of Artificial Intelligence · Second affiliation of several PKU authors.115

Released
December 20241718
More

arXiv v1 on 2024-12-18; Hugging Face dataset created 2025-01-02

Version
v1.2. RoboMIND 2.0 is a separate dataset.21719
More

Current release v1.2 (dataset card). Paper: arXiv v1 2024-12-18, v2 2025-02-14, v3 2025-05-27 (RSS 2025 version). Release folders benchmark1_0, benchmark1_1, benchmark1_2.

v1.0 · arXiv v1 (2024-12): 55k trajectories, 279 tasks, 61 object classes, 36 skills, 308.6 hours, 4 robot types. The card's Version 1.0 note says 69 object classes.122

v1.1 · 107k trajectories, 479 tasks, 96 object classes (card). Matches arXiv v2 (2025-02) and v3.281

v1.2 · Adds 10 'Upright_Cup' tasks: 1 real-world task and 9 from the digital twin, varying mug placement, table texture and mug appearance. Release date not stated.219

RoboMIND 2.0 (separate dataset) · Same lead organisation, arXiv 2025-12-31: over 310K dual-arm trajectories on six robot types, mobile and tactile data, and an Isaac Sim benchmark (RoboMIND-Sim). Its paper calls the earlier set 'RoboMIND 1.0' and says 1.0 focuses on single-arm manipulation. The v1 card and site announce it as 'RoboMIND V2.0'. Separate Atlas record (robomind-2-0).9211+1

Last update
The data last changed in April 2026.18621+1
More

Hugging Face files last changed 2026-04-14, the day the data-check repository (RoboMIND-dataset-utils) was created. ModelScope copy last updated 2026-07-20. Training toolchain last pushed 2026-09-16.

Status
Maintained. New data collection goes into RoboMIND 2.0. Inferred1869
More

Data checks and fixes in 2026-04; new collection goes into RoboMIND 2.0

Five open user discussions on the Hugging Face card (2025-12 to 2026-08) have no maintainer reply visible.

Setup

Runs in
Real robots1
More

Headline results come from real robots. The paper also evaluates policies in its Isaac Sim digital twin on 6 tasks (see sim_to_real and validity). The v1 twin environment is not released; only its recorded trajectories are.

Simulator
Isaac Sim, used only for the digital twin (a simulated copy of the real set-up)1222
More

Used for the twin data and the twin evaluations. In the released simulation data, robot state is sampled at about four times the camera rate, and depth images are not yet available (card). The Isaac Sim code released later (RoboMIND-Sim, Isaac Sim 4.5 and 5.1) is for RoboMIND 2.0 Tien Kung tasks.

Robot
One arm, Two arms, Humanoid, Robot hand1
Robot model
4 robot types1
More

Franka Emika Panda (3 RealSense D435i cameras, Robotiq gripper); UR5e (top RealSense camera, Robotiq gripper); AgileX Cobot Magic V2.0 dual-arm (3 Orbbec cameras); X-Humanoid Tien Kung humanoid (42 DoF, two Inspire RH56DFX dexterous hands, Orbbec Gemini 335 cameras)

The v1.1 release also has folders for a dual-arm Franka FR3 setup and a second Tien Kung variant (h5_franka_fr3_dual, h5_tienkung_prod1_gello_1rgb), which the paper does not describe.

Setting
Lab tables themed as kitchen, home or office7
More

Lab setups themed as kitchen 43.4%, domestic 26.7%, office 16.5%, retail 6.9%, industrial 6.5% of trajectories

Shares read from Figure 1(d) of the RSS version. The authors list 'relatively simple background environments' as a limitation.

Tasks
479 tasks1212+1
More

479 tasks (v1.1 and v1.2); 279 in v1.0

Tasks are defined by robot, skill, objects and scene. The v1.2 instruction file lists 479 rows with 468 distinct task names (counted by us).

Training data
107k trajectories, about 28% of them simulated1211
More

107k trajectories (305.5 hours) on four robot types. By the paper's own breakdown, 30,035 of them (about 28%) come from the simulated digital twin, although the site and card call all 107k 'real-world'.

Per-robot counts differ between the paper text and its own Figure 1 / the card; resolved as far as possible in the items. See issues.i1 and issues.i2.

Paper text: real per robot plus a simulation total · Franka 26,856; Tien Kung 15,187; AgileX 10,269; UR5e 25,170; simulation 30,035. Sum 107,517 (our arithmetic).17

Figure 1(a) and dataset card: totals per robot including simulation · Franka 52,926; humanoid (Tien Kung) 19,152; AgileX 10,629; UR5e 25,170. Sum 107,877 (our arithmetic). The card calls all of them teleoperation data.72

How the two fit · Franka 52,926 = 26,856 real + 26,070 twin trajectories (Section IV-A gives 'over 26,070' twin and 26,866 real). Tien Kung 19,152 = 15,187 real + 3,965 twin, by our arithmetic (30,035 − 26,070); the paper's 17.8% humanoid share and its '19k' humanoid ablation fit this. The v1.1 release has a Tien Kung simulation folder, although Section IV-A says the other three robots have real data only. AgileX 10,629 (figure, card) vs 10,269 (text) is unresolved; the 360 difference is the whole gap between the two totals.172+2

v1.0 breakdown · Franka 19,222; Tien Kung 9,686; AgileX 8,030; UR-5e 6,911; simulation 11,783 (all Franka). Sum 55,632 (our arithmetic). 308.6 hours.12

Average length · Franka 179 frames, UR 158, AgileX 655, humanoid 669 (Figure 1(b))7

Failures and annotations · 5k failure demonstrations with causes; 10k trajectories with frame-level language annotations (Gemini drafts revised by people)1

12.3 TB · Total file size on Hugging Face2

Changes at test
Only one task is tested with new objects and backgrounds.1
More

Most tests reuse the training tasks. One task was also tested with new objects and new tablecloths.

Table VI: FR-PlaceBreadPlate with the bread swapped for corn, banana or apple, and with three unseen tablecloths.

Scoring and access

Scored by
Success rate1
More

Testers record success or failure, and the cause of each failure from 9 predefined categories.

Score
Success rate over 10 real-robot trials per task17
More

Each trained model is run 10 times per task on the real robot and the share of successes is reported, with failure causes logged.

Single-task imitation learning · ACT, Diffusion Policy and BAKU trained from scratch on 45 tasks (Franka 15, Tien Kung 10, AgileX 15, UR5e 5). ACT averages: AgileX 55.3%, UR5e 38.0%, Tien Kung 34.0%, Franka 30.7%.1

Vision-language-action models · OpenVLA (Franka only), RDT-1B and CrossFormer fine-tuned per robot on 15 tasks. Example: RDT-1B 10/10 on AX-AppleYellowPlate; CrossFormer 0/10 on all five AgileX tasks.1

Pretraining on all of RoboMIND · RDT-1B and CrossFormer pretrained on the full set, then fine-tuned on about 1% of it, improved on most of the 15 tasks (Table IV). Excluding the humanoid data (19k of 107k) lowered RDT-1B's average on 5 Franka tasks from 0.68 to 0.6.1

Most common failure · For ACT, 'Inaccurate Positioning' is the top failure cause on every robot type (48% of failures on the humanoid).1

Trials
10 per task1
More

10 per task and model

Who runs it
Each team tests its own model Inferred1
More

All results come from the authors' own tests.

Error bars
Not reported Inferred1
More

Results are given as successes out of 10 or as percentages, with no error bars or intervals.

Leaderboard
None. Scores are only in papers. Inferred112
More

No leaderboard, test server or challenge on the project site, dataset card or website repo.

Code licence
Apache-2.0 for the training toolchain the paper links (x-humanoid-training-toolchain) and for the data-check scripts (RoboMIND-dataset-utils)56
More

Toolchain LICENSE: Apache 2.0, 'Copyright 2025 The Beijing Innovation Center of Humanoid Robotics'. The project-website repo has an Apache 2.0 LICENSE file, which GitHub reports as NOASSERTION.

Data licence
Apache-2.01821
More

Hugging Face card metadata 'apache-2.0'; ModelScope copy 'apache-2.0'.

Asset licence
Unknown
More

The digital-twin scenes and 3D assets for version 1 are not released, and no licence for them is stated. Looked at the paper, dataset card, project site, ModelScope file tree and the X-Humanoid GitHub organisation.

Access
Free. Users agree to share contact details and are approved automatically.21821+1
More

Hugging Face gate with automatic approval (agree to share contact information). Also on ModelScope and BAAI's data platform.

Hub API: gated = 'auto'. The ModelScope file tree is publicly listable; its download conditions (ApprovalMode 1, ProtectedMode 2) were not tested.

Commercial use
Allowed Inferred185
More

Data and code are Apache-2.0, which allows commercial use. No third-party assets are distributed. Not legal advice.

Published at
Robotics: Science and Systems XXI (Los Angeles, June 21-25, 2025), paper 152, DOI 10.15607/RSS.2025.XXI.15215724

Sources 24

  1. 1RoboMIND full text v3 (RSS 2025 version)Paper · May 2025 · checked 10 Oct 2026
  2. 2x-humanoid-robomind/RoboMIND dataset card (composition, version notes, data notes)Repository · 14 Apr 2026 · checked 10 Oct 2026
  3. 3EO-1: An Open Unified Embodied Foundation Model for General Robot Control (training data table)Paper · Aug 2025 · checked 10 Oct 2026
  4. 4XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations (pretraining data)Paper · Nov 2025 · checked 10 Oct 2026
  5. 5Open-X-Humanoid/x-humanoid-training-toolchain (README, LICENSE, GitHub API record)Repository · 16 Sep 2026 · checked 10 Oct 2026
  6. 6Open-X-Humanoid/RoboMIND-dataset-utils README and LICENSERepository · 14 Apr 2026 · checked 10 Oct 2026
  7. 7RoboMIND, RSS 2025 PDF (Figure 1, Table VII, Figure 17)Paper · Jun 2025 · checked 10 Oct 2026
  8. 8RoboMIND full text v2Paper · Feb 2025 · checked 10 Oct 2026
  9. 9RoboMIND 2.0 full text v3 (introduction and Table 1)Paper · Feb 2026 · checked 10 Oct 2026
  10. 10加速打造全球具身智能数据平台 北京人形开源数据集下载量破2000万Secondary · 2 Sep 2026 · checked 10 Oct 2026
  11. 11RoboMIND project siteOfficial site · Jun 2026 · checked 10 Oct 2026
  12. 12RoboMIND full text v1Paper · Dec 2024 · checked 10 Oct 2026
  13. 13RoboMIND discussion #6: UR5e data with almost no movement (user reports, open)Repository · Jan 2026 · checked 10 Oct 2026
  14. 14RoboMIND discussion #8: invalid end-effector poses in h5_ur_1rgb (user report, open)Repository · Mar 2026 · checked 10 Oct 2026
  15. 15RoboMIND, Robotics: Science and Systems XXI proceedings page (paper 152)Paper · Jun 2025 · checked 10 Oct 2026
  16. 16北京人形机器人创新中心有限公司 official site (home page)Official site · 2026 · checked 10 Oct 2026
  17. 17RoboMIND (arXiv abstract page, submission history v1 to v3)Paper · Dec 2024 · checked 10 Oct 2026
  18. 18Hugging Face Hub API: x-humanoid-robomind/RoboMIND (licence, gated, downloads, lastModified)Index · 10 Oct 2026 · checked 10 Oct 2026
  19. 19ModelScope file tree: X-Humanoid/RoboMIND (benchmark1_0, 1_1, 1_2 folders)Repository · 2026 · checked 10 Oct 2026
  20. 20Open-X-Humanoid/RoboMIND-Sim README (RoboMIND 2.0 Isaac Sim benchmark)Repository · Mar 2026 · checked 10 Oct 2026
  21. 21ModelScope API: X-Humanoid/RoboMIND (licence, downloads, dates)Index · 20 Jul 2026 · checked 10 Oct 2026
  22. 22RoboMIND all_robot_h5_info_v1.2.md (file formats incl. Simulation Franka and Simulation Tien Kung)Repository · 2025 · checked 10 Oct 2026
  23. 23RoboMIND_v1_2_instr.csv (task instruction list)Repository · 2025 · checked 10 Oct 2026
  24. 24Semantic Scholar API record for arXiv:2412.13877Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information

validity: arXiv v1, v2 and v3 full text (the correlation table first appears in v3; the co-training figure in v2), RSS 2025 PDF pages 1-2 and 11-15, project site, dataset card, ModelScope card and file tree, X-Humanoid GitHub organisation (RoboMIND-Sim and x-humanoid-vla-simulation-benchmark belong to RoboMIND 2.0 or other projects). No independent sim-versus-real study of RoboMIND was found; papers that use the data train on it and do not report RoboMIND scores.

count conflict (per-robot trajectories): All three arXiv versions, RSS PDF Figure 1, Hugging Face and ModelScope cards, ModelScope folder listing for benchmark1_0, 1_1 and 1_2, all_robot_h5_info_v1.2.md, RoboMIND_v1_2_instr.csv. Archives are split tar.gz parts, so per-folder episode counts were not verified.

license_code and license_assets: Paper link to x-humanoid-training-toolchain (redirects to Open-X-Humanoid), its LICENSE file, RoboMIND-dataset-utils LICENSE, website repo LICENSE, dataset card, ModelScope metadata.

official Chinese announcement: X-Humanoid official site home and news pages (the open-source portal is script-rendered); Beijing News report of 2026-09-02 (secondary). The shared web-search budget ran out before an official RoboMIND v1 launch article was found.

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
  2. Expanded to a full entry. Resolved the per-robot count conflict at the source as far as possible (paper text vs Figure 1 and card; twin data folded into robot totals; one AgileX figure unresolved). Added the code licence (Apache-2.0 toolchain), builder legal form, data-quality findings from X-Humanoid's own check scripts, two author-run sim-versus-real comparisons, and the relation to RoboMIND 2.0.