RoboMIND
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
RoboMIND is a dataset of about 107,000 demonstrations on four robot types, used to train robot control models (policies). About 28% of the demonstrations come from simulation, and its paper tests policies on real robots.12
What a score here does not tell you Inferred
- Whether a policy would get the same score in another lab.All tests in the paper ran on X-Humanoid's own robot setups.
- How well a policy does in new surroundings.In the one task tested this way, new tablecloths cut success to 0 to 2 of 10 trials.
- Whether the digital twin (the simulated copy) predicts real-robot results.The authors checked it with only 2 policies on 5 tasks.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| RoboMIND digital-twin check (Table VII) May 2025 | Pearson correlation r = 0.83 (ACT) and 0.91 (Diffusion Policy), each across 5 tasks17 | ACT and Diffusion Policy, trained on real data, were each run 10 times on the same 5 Franka tasks in the Isaac Sim twin and on the real robot. The study’s authors described the result as “positive correlations”. | The benchmark’s authors |
| RoboMIND co-training study (Figure 17) Feb 2025 | No statistic was reported. The policy trained only on twin data scored 0.9 in the twin and 0.1 on the real robot. The policy trained only on real data scored 0 in the twin and 0.6 on the real robot.87 | 7 ACT policies trained on different mixes of real and twin data were each scored on one task (FR-UprightBlueCup) in the twin and on the real robot. The study’s authors described the result as “simulation data alone is insufficient”. | The benchmark’s authors |
Our assessment Opinion
RoboMIND is mainly training data. Its test results are one-off tests by its builders.
Reasoning
RoboMIND is mainly training data. Its paper's success rates are one-off tests by the builders on their own robots and task setups, with 10 trials each and no error bars. They are not comparable with results in other papers.
Confidence: high
The check of the digital twin is too small to show that the twin can be used as a test.
Reasoning
The digital-twin check is small. Two policies on five tasks, with ten trials each, cannot show that the twin ranks many policies correctly. The twin also scored policies far lower than the real robot. A policy trained only in the twin scored well there and failed on the real robot.
Confidence: high
When quoting its size, name the version and say whether simulated data is counted.
Reasoning
When quoting its size, name the version and say whether simulation is included. Version 1.0 has 55k trajectories, and versions 1.1 and 1.2 have 107k, of which about 30k are simulated. RoboMIND 2.0, a different dataset, has over 310K.
Confidence: high
Known problems 6
The name is shared with RoboMIND 2.0
The name 'RoboMIND' is often used for RoboMIND 2.0, a separate and larger dataset. Their sizes and download counts are easy to mix up.2910
Details
The version 1 card and site announce 'RoboMIND V2.0', a different and larger dataset (over 310K trajectories, six robot types). Chinese news coverage uses 'RoboMIND' for 2.0 figures (300,000+ trajectories, 700+ tasks, over 20 million downloads). Sizes and download counts are easy to mix up.
About 28% of the 'real-world' data is simulated
The site and dataset card describe '107k real-world' trajectories. By the paper's own breakdown, about 30,000 of them come from simulation.1121+1
Details
The project site and dataset card describe '107k real-world demonstration trajectories'. The paper's own breakdown counts 30,035 of them from the Isaac Sim digital twin, and Franka's 49.2% share includes over 26,070 twin trajectories. The paper's comparison table places RoboMIND among real-world datasets. Version 1.0 also mixed in 11,783 simulated trajectories while its abstract said '55k real-world'.
Trajectory counts differ between versions, the paper's figure and its text
The number of trajectories differs between versions, between the paper's text and its figure, and on the dataset card. Some of these differences are still unexplained.1217+1
Details
Version 1.0 (arXiv v1) has 55k trajectories and 308.6 hours; versions 1.1 and 1.2 have 107k trajectories but 305.5 hours. Within the RSS paper, the text lists Franka 26,856, Tien Kung 15,187, AgileX 10,269, UR5e 25,170 and simulation 30,035, while Figure 1(a) and the card list Franka 52,926, humanoid 19,152, AgileX 10,629 and UR5e 25,170. Most of the gap is explained by Figure 1 and the card folding twin data into each robot (see facts.demonstrations). Still unresolved: AgileX 10,629 vs 10,269; Franka real 26,856 (introduction) vs 26,866 (Section IV-A); object classes 61 vs 69 for v1.0; and the drop in hours from 308.6 to 305.5 while trajectories doubled.
End-effector data is frozen in some subsets
X-Humanoid's own checks found that end-effector values (the position and angle of the robot's gripper or hand) stay almost constant in three subsets. Users report further problems.6213+1
Details
X-Humanoid's data-check repository (2026-04) reports end-effector values that stay almost constant in three subsets: h5_ur_1rgb, h5_simulation and h5_sim_franka_3rgb, probably from a blocking problem during collection. It recommends using joint positions and recomputing end-effector poses with forward kinematics. Users report constant UR end-effector poses (discussion #8) and UR episodes whose joints barely move although the video shows motion (discussion #6); one user posted a scripted count of 22,987 of 25,721 UR episodes (89.37%) with the first six joint values unchanged, which X-Humanoid has not answered. The card also notes BGR image order in four subsets and RGB in the rest, and 675 Franka trajectories with only two of three cameras.
The digital twin check is small and cannot be repeated
The digital twin was checked with 2 policies on 5 tasks. The twin is not released, so others cannot repeat the check.17
Details
Table VII compares 2 policies on 5 Franka tasks with 10 trials per task; each Pearson r rests on 5 points and has no interval. The twin scored policies far lower than the real robot (ACT 1.4/10 vs 5.4/10 on average). In Figure 17 a policy trained only on twin data scored 0.9 in the twin and 0.1 on the real robot, which the authors attribute to physics gaps in contact-rich tasks. The version 1 twin is not released, so others cannot rerun the check.
The backgrounds are simple and fixed
In the authors' test, three new tablecloths cut success on one task to 0 to 2 of 10 trials.1
Details
The authors list relatively simple background environments as a limitation. In their generalisation test, three unseen tablecloths cut FR-PlaceBreadPlate success to 0-2 of 10 for OpenVLA, RDT-1B and CrossFormer (from 4, 9 and 10 of 10).
Details
About
- What it is
- Dataset Inferred12
More
The title says 'Benchmark', and the paper says RoboMIND 'serves as a benchmark' for the methods it tests. What is released is training data plus conversion and training code; the real-robot test setups are not released.
- Built by
- X-Humanoid, with Peking University and BAAI11511
More
37 authors on the RSS page.
Beijing Innovation Center of Humanoid Robotics (X-Humanoid) · 北京人形机器人创新中心有限公司. Builds the Tien Kung humanoid. Corresponding author Jian Tang; project leaders Zhengping Che and Xiaozhu Ju.116
Peking University · State Key Laboratory of Multimedia Information Processing, School of Computer Science. Corresponding author Shanghang Zhang.1
Beijing Academy of Artificial Intelligence · Second affiliation of several PKU authors.115
- Version
- v1.2. RoboMIND 2.0 is a separate dataset.21719
More
Current release v1.2 (dataset card). Paper: arXiv v1 2024-12-18, v2 2025-02-14, v3 2025-05-27 (RSS 2025 version). Release folders benchmark1_0, benchmark1_1, benchmark1_2.
v1.0 · arXiv v1 (2024-12): 55k trajectories, 279 tasks, 61 object classes, 36 skills, 308.6 hours, 4 robot types. The card's Version 1.0 note says 69 object classes.122
v1.1 · 107k trajectories, 479 tasks, 96 object classes (card). Matches arXiv v2 (2025-02) and v3.281
v1.2 · Adds 10 'Upright_Cup' tasks: 1 real-world task and 9 from the digital twin, varying mug placement, table texture and mug appearance. Release date not stated.219
RoboMIND 2.0 (separate dataset) · Same lead organisation, arXiv 2025-12-31: over 310K dual-arm trajectories on six robot types, mobile and tactile data, and an Isaac Sim benchmark (RoboMIND-Sim). Its paper calls the earlier set 'RoboMIND 1.0' and says 1.0 focuses on single-arm manipulation. The v1 card and site announce it as 'RoboMIND V2.0'. Separate Atlas record (robomind-2-0).9211+1
Setup
- Runs in
- Real robots1
More
Headline results come from real robots. The paper also evaluates policies in its Isaac Sim digital twin on 6 tasks (see sim_to_real and validity). The v1 twin environment is not released; only its recorded trajectories are.
- Simulator
- Isaac Sim, used only for the digital twin (a simulated copy of the real set-up)1222
More
Used for the twin data and the twin evaluations. In the released simulation data, robot state is sampled at about four times the camera rate, and depth images are not yet available (card). The Isaac Sim code released later (RoboMIND-Sim, Isaac Sim 4.5 and 5.1) is for RoboMIND 2.0 Tien Kung tasks.
- Robot
- One arm, Two arms, Humanoid, Robot hand1
- Robot model
- 4 robot types1
More
Franka Emika Panda (3 RealSense D435i cameras, Robotiq gripper); UR5e (top RealSense camera, Robotiq gripper); AgileX Cobot Magic V2.0 dual-arm (3 Orbbec cameras); X-Humanoid Tien Kung humanoid (42 DoF, two Inspire RH56DFX dexterous hands, Orbbec Gemini 335 cameras)
The v1.1 release also has folders for a dual-arm Franka FR3 setup and a second Tien Kung variant (h5_franka_fr3_dual, h5_tienkung_prod1_gello_1rgb), which the paper does not describe.
- Setting
- Lab tables themed as kitchen, home or office7
More
Lab setups themed as kitchen 43.4%, domestic 26.7%, office 16.5%, retail 6.9%, industrial 6.5% of trajectories
Shares read from Figure 1(d) of the RSS version. The authors list 'relatively simple background environments' as a limitation.
- Tasks
- 479 tasks1212+1
More
479 tasks (v1.1 and v1.2); 279 in v1.0
Tasks are defined by robot, skill, objects and scene. The v1.2 instruction file lists 479 rows with 468 distinct task names (counted by us).
- Training data
- 107k trajectories, about 28% of them simulated1211
More
107k trajectories (305.5 hours) on four robot types. By the paper's own breakdown, 30,035 of them (about 28%) come from the simulated digital twin, although the site and card call all 107k 'real-world'.
Per-robot counts differ between the paper text and its own Figure 1 / the card; resolved as far as possible in the items. See issues.i1 and issues.i2.
Paper text: real per robot plus a simulation total · Franka 26,856; Tien Kung 15,187; AgileX 10,269; UR5e 25,170; simulation 30,035. Sum 107,517 (our arithmetic).17
Figure 1(a) and dataset card: totals per robot including simulation · Franka 52,926; humanoid (Tien Kung) 19,152; AgileX 10,629; UR5e 25,170. Sum 107,877 (our arithmetic). The card calls all of them teleoperation data.72
How the two fit · Franka 52,926 = 26,856 real + 26,070 twin trajectories (Section IV-A gives 'over 26,070' twin and 26,866 real). Tien Kung 19,152 = 15,187 real + 3,965 twin, by our arithmetic (30,035 − 26,070); the paper's 17.8% humanoid share and its '19k' humanoid ablation fit this. The v1.1 release has a Tien Kung simulation folder, although Section IV-A says the other three robots have real data only. AgileX 10,629 (figure, card) vs 10,269 (text) is unresolved; the 360 difference is the whole gap between the two totals.172+2
v1.0 breakdown · Franka 19,222; Tien Kung 9,686; AgileX 8,030; UR-5e 6,911; simulation 11,783 (all Franka). Sum 55,632 (our arithmetic). 308.6 hours.12
Average length · Franka 179 frames, UR 158, AgileX 655, humanoid 669 (Figure 1(b))7
Failures and annotations · 5k failure demonstrations with causes; 10k trajectories with frame-level language annotations (Gemini drafts revised by people)1
12.3 TB · Total file size on Hugging Face2
- Changes at test
- Only one task is tested with new objects and backgrounds.1
More
Most tests reuse the training tasks. One task was also tested with new objects and new tablecloths.
Table VI: FR-PlaceBreadPlate with the bread swapped for corn, banana or apple, and with three unseen tablecloths.
Scoring and access
- Scored by
- Success rate1
More
Testers record success or failure, and the cause of each failure from 9 predefined categories.
- Score
- Success rate over 10 real-robot trials per task17
More
Each trained model is run 10 times per task on the real robot and the share of successes is reported, with failure causes logged.
Single-task imitation learning · ACT, Diffusion Policy and BAKU trained from scratch on 45 tasks (Franka 15, Tien Kung 10, AgileX 15, UR5e 5). ACT averages: AgileX 55.3%, UR5e 38.0%, Tien Kung 34.0%, Franka 30.7%.1
Vision-language-action models · OpenVLA (Franka only), RDT-1B and CrossFormer fine-tuned per robot on 15 tasks. Example: RDT-1B 10/10 on AX-AppleYellowPlate; CrossFormer 0/10 on all five AgileX tasks.1
Pretraining on all of RoboMIND · RDT-1B and CrossFormer pretrained on the full set, then fine-tuned on about 1% of it, improved on most of the 15 tasks (Table IV). Excluding the humanoid data (19k of 107k) lowered RDT-1B's average on 5 Franka tasks from 0.68 to 0.6.1
Most common failure · For ACT, 'Inaccurate Positioning' is the top failure cause on every robot type (48% of failures on the humanoid).1
- Trials
- 10 per task1
More
10 per task and model
- Who runs it
- Each team tests its own model Inferred1
More
All results come from the authors' own tests.
- Error bars
- Not reported Inferred1
More
Results are given as successes out of 10 or as percentages, with no error bars or intervals.
- Leaderboard
- None. Scores are only in papers. Inferred112
More
No leaderboard, test server or challenge on the project site, dataset card or website repo.
- Code licence
- Apache-2.0 for the training toolchain the paper links (x-humanoid-training-toolchain) and for the data-check scripts (RoboMIND-dataset-utils)56
More
Toolchain LICENSE: Apache 2.0, 'Copyright 2025 The Beijing Innovation Center of Humanoid Robotics'. The project-website repo has an Apache 2.0 LICENSE file, which GitHub reports as NOASSERTION.
- Data licence
- Apache-2.01821
More
Hugging Face card metadata 'apache-2.0'; ModelScope copy 'apache-2.0'.
- Asset licence
- Unknown
More
The digital-twin scenes and 3D assets for version 1 are not released, and no licence for them is stated. Looked at the paper, dataset card, project site, ModelScope file tree and the X-Humanoid GitHub organisation.
- Access
- Free. Users agree to share contact details and are approved automatically.21821+1
More
Hugging Face gate with automatic approval (agree to share contact information). Also on ModelScope and BAAI's data platform.
Hub API: gated = 'auto'. The ModelScope file tree is publicly listable; its download conditions (ApprovalMode 1, ProtectedMode 2) were not tested.
Sources 24
- 1RoboMIND full text v3 (RSS 2025 version)Paper · May 2025 · checked 10 Oct 2026
- 2x-humanoid-robomind/RoboMIND dataset card (composition, version notes, data notes)Repository · 14 Apr 2026 · checked 10 Oct 2026
- 3EO-1: An Open Unified Embodied Foundation Model for General Robot Control (training data table)Paper · Aug 2025 · checked 10 Oct 2026
- 4XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations (pretraining data)Paper · Nov 2025 · checked 10 Oct 2026
- 5Open-X-Humanoid/x-humanoid-training-toolchain (README, LICENSE, GitHub API record)Repository · 16 Sep 2026 · checked 10 Oct 2026
- 6Open-X-Humanoid/RoboMIND-dataset-utils README and LICENSERepository · 14 Apr 2026 · checked 10 Oct 2026
- 7RoboMIND, RSS 2025 PDF (Figure 1, Table VII, Figure 17)Paper · Jun 2025 · checked 10 Oct 2026
- 8RoboMIND full text v2Paper · Feb 2025 · checked 10 Oct 2026
- 9RoboMIND 2.0 full text v3 (introduction and Table 1)Paper · Feb 2026 · checked 10 Oct 2026
- 10加速打造全球具身智能数据平台 北京人形开源数据集下载量破2000万Secondary · 2 Sep 2026 · checked 10 Oct 2026
- 11RoboMIND project siteOfficial site · Jun 2026 · checked 10 Oct 2026
- 12RoboMIND full text v1Paper · Dec 2024 · checked 10 Oct 2026
- 13RoboMIND discussion #6: UR5e data with almost no movement (user reports, open)Repository · Jan 2026 · checked 10 Oct 2026
- 14RoboMIND discussion #8: invalid end-effector poses in h5_ur_1rgb (user report, open)Repository · Mar 2026 · checked 10 Oct 2026
- 15RoboMIND, Robotics: Science and Systems XXI proceedings page (paper 152)Paper · Jun 2025 · checked 10 Oct 2026
- 16北京人形机器人创新中心有限公司 official site (home page)Official site · 2026 · checked 10 Oct 2026
- 17RoboMIND (arXiv abstract page, submission history v1 to v3)Paper · Dec 2024 · checked 10 Oct 2026
- 18Hugging Face Hub API: x-humanoid-robomind/RoboMIND (licence, gated, downloads, lastModified)Index · 10 Oct 2026 · checked 10 Oct 2026
- 19ModelScope file tree: X-Humanoid/RoboMIND (benchmark1_0, 1_1, 1_2 folders)Repository · 2026 · checked 10 Oct 2026
- 20Open-X-Humanoid/RoboMIND-Sim README (RoboMIND 2.0 Isaac Sim benchmark)Repository · Mar 2026 · checked 10 Oct 2026
- 21ModelScope API: X-Humanoid/RoboMIND (licence, downloads, dates)Index · 20 Jul 2026 · checked 10 Oct 2026
- 22RoboMIND all_robot_h5_info_v1.2.md (file formats incl. Simulation Franka and Simulation Tien Kung)Repository · 2025 · checked 10 Oct 2026
- 23RoboMIND_v1_2_instr.csv (task instruction list)Repository · 2025 · checked 10 Oct 2026
- 24Semantic Scholar API record for arXiv:2412.13877Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
validity: arXiv v1, v2 and v3 full text (the correlation table first appears in v3; the co-training figure in v2), RSS 2025 PDF pages 1-2 and 11-15, project site, dataset card, ModelScope card and file tree, X-Humanoid GitHub organisation (RoboMIND-Sim and x-humanoid-vla-simulation-benchmark belong to RoboMIND 2.0 or other projects). No independent sim-versus-real study of RoboMIND was found; papers that use the data train on it and do not report RoboMIND scores.
count conflict (per-robot trajectories): All three arXiv versions, RSS PDF Figure 1, Hugging Face and ModelScope cards, ModelScope folder listing for benchmark1_0, 1_1 and 1_2, all_robot_h5_info_v1.2.md, RoboMIND_v1_2_instr.csv. Archives are split tar.gz parts, so per-folder episode counts were not verified.
license_code and license_assets: Paper link to x-humanoid-training-toolchain (redirects to Open-X-Humanoid), its LICENSE file, RoboMIND-dataset-utils LICENSE, website repo LICENSE, dataset card, ModelScope metadata.
official Chinese announcement: X-Humanoid official site home and news pages (the open-source portal is script-rendered); Beijing News report of 2026-09-02 (secondary). The shared web-search budget ran out before an official RoboMIND v1 launch article was found.
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
- Expanded to a full entry. Resolved the per-robot count conflict at the source as far as possible (paper text vs Figure 1 and card; twin data folded into robot totals; one AgileX figure unresolved). Added the code licence (Apache-2.0 toolchain), builder legal form, data-quality findings from X-Humanoid's own check scripts, two author-run sim-versus-real comparisons, and the relation to RoboMIND 2.0.