RoboTwin 2.0
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
RoboTwin 2.0 is a simulated benchmark and data generator for two-armed robots. It tests robot policies (the models that control a robot) on 50 tabletop tasks in clean and randomized scenes, and it has an official leaderboard.123
What a score here does not tell you Inferred
- How well a policy will do on a real robot.No study has scored the same policies on RoboTwin 2.0 and on real robots.
- How a score compares with scores in other papers.Papers use two different training protocols, and they give very different Hard scores.
- How well a policy works on other robots.All scores use one simulated robot, the Aloha-AgileX.
- Whether the task stays done after a success is recorded.A success is recorded the first moment the goal is met, even if an object is still moving or falling.
Comparisons with real robots
Tried on real robots, not compared Inferred189+1
Details
Data-transfer evidence only. The paper trained RDT on a real COBOT-Magic robot for 4 tasks: with 10 real demonstrations plus 1,000 RoboTwin trajectories, average success on unseen cluttered backgrounds rose from 9.0% to 42.0%; with synthetic data only it reached 29.5% (arXiv: 367% and 228% relative gains; ICML abstract: '3.6x' and '2.2x'). Trial counts for these real runs are not stated. Independent check on the earlier version only: WorldEval (2025-05) ran real-trained policies in RoboTwin 1.0 on 3 custom tasks, with images translated by MidJourney, and found average Pearson r 0.411 and MMRV 0.261 against real results. That study predates 2.0 and uses other tasks, so it is not counted as validity for 2.0. A May 2026 paper by other authors also marks RoboTwin as having no validated sim-to-real correlation in its comparison table.
Our assessment Opinion
RoboTwin 2.0 is a useful simulation test. Its link to real-robot results has not been measured.
Reasoning
RoboTwin 2.0 tests two-arm skills under visual clutter, and the 2026 audit found no cheap shortcut on it. A high score still says little about real robots. No study has scored the same policies on version 2.0 and on real hardware.
Confidence: medium
Check which training protocol a score used.
Reasoning
Before comparing numbers, check the training protocol. The official leaderboard uses 50 clean demonstrations per task. The other protocol adds 500 randomized demonstrations per task. Hard scores under the two are not comparable.
Confidence: high
Look at the Hard scores as well as the average.
Reasoning
Read Easy and Hard separately. Several policies do well in clean scenes and score near zero in randomized ones. On the leaderboard, FastWAM scores 77.8 on Easy and 1.9 on Hard. The leaderboard's average hides this.
Confidence: high
Ignore small gaps between leaderboard entries.
Reasoning
Small gaps between leaderboard entries are hard to interpret. There are no error bars, and entries were produced by different teams on different code and data versions. Some entries appear to use 50 rather than 100 trials.
Confidence: medium
Known problems 9
Two training protocols give very different Hard scores
Depending on the training data, Hard scores for the same model range from about 2% to 92%.2311+2
Details
The official board trains on 50 clean demonstrations per task; its best Hard score is 68.12. Since Motus (2025-12) many papers train one policy on 50 clean plus 500 randomized demonstrations per task, which puts randomized scenes into training; MotuBrain reports 96.1 Hard. The same model names differ widely: Fast-WAM 91.78 Hard in its paper vs FastWAM 1.9 Hard on the board; the audit's data file lists starVLA-α at 88.3 and ABot-M0 at 85.08, while the board shows starVLA 3.16 and Abot-M0 30.36. Both protocols are called 'RoboTwin 2.0' in papers.
The same baseline gets different scores in different sources
Different sources report the Hard score of π0.5 as 46.0, 43.84 and 76.76.2311+2
Details
π0.5: 70.7 / 46.0 on the official board; 42.98 / 43.84 in Motus's run and 82.74 / 76.76 in Fast-WAM's table, both under the 50+500 protocol. X-VLA: 70.0 / 39.0 in its own paper (training data not stated); 68.0 / 20.9 on the board; 72.8 / 72.84 in Motus's run. The audit replaced X-VLA's own number with a protocol-matched one because X-VLA did not train on 50 clean plus 500 randomized demonstrations.
Code and data changed while the leaderboard was running
Leaderboard entries were produced with different versions of the code and data.141516+2
Details
DP3 evaluation code was fixed on 2025-07-19 and ACT deployment code on 2025-08-25, with the board updated. Several success checks changed in July and August 2025. The official demo_clean training data was refreshed on 2026-09-14 with 'legal-joint replacements', after most co-train entries were listed. Paper v1 reported clean-to-randomized drops of 28.2 (RDT) and 40.3 (π0) points on 13 tasks; v2 reports 20.8 and 30.1 on 50 tasks. Board entries date from 2025-08 to 2026-10 and so span these changes.
Seeds did not reproduce the same scene until September 2026
The same seed (the number that sets up a random scene) could produce a different scene. This was fixed in September 2026. A separate bug in batch evaluation is still open.171819
Details
Users found that a dataset seed did not recreate the recorded scene (issue #288). Causes: asset order depended on the file system in 5 tasks, table-clutter bounds grew with each setup call, and Python's random generator, which picks the instruction, was not seeded. Fixes were merged on 2026-09-24 (PRs #490 and #509). A separate open bug (#522, 2026-10-09) shows batch evaluation can score a different set of seeds depending on the number of workers.
Success checks can count brief moments and some disagree with the instructions
A success can be recorded while an object is still falling. Some instructions ask for something different from what the success check tests.202122+3
Details
An episode ends as a success the first time the task predicate holds, without checking that objects stay put; users report bottles and stacked blocks counted as successes while still moving or falling (issues #493 and #491, open). rotate_qrcode instructions asked to lift and rotate while the check required putting the object down (fixed 2026-09-22); some hanging_mug instructions name the wrong arm (#521, open). LeRobot's docs say open_laptop's check fails during normal policy evaluation; RoboTwin fixed the arm-tag setup for open_laptop, place_object_scale and put_object_cabinet on 2026-08-20 (PR #488), after LeRobot's note.
Image colour channels differ between data versions
Older data stores images with the red and blue channels swapped. The official decoder handles both formats.142615
Details
The README says older ('legacy') data stores red and blue swapped and was never migrated, so old and new formats can appear in one training run; images must be decoded with decode_image_bit. A user reported the swap in 2026-04 (issue #438); a maintainer replied that images follow RGB order. The demo_clean data and LeRobot archives were refreshed with standard RGB images on 2026-09-14.
Some leaderboard results cannot be re-checked
The trained models (checkpoints) behind several leaderboard baselines are not available. Users report gaps when they try to reproduce some results.271528+2
Details
The π0 checkpoints behind the board were never released; a maintainer said π0 was not part of the latest re-evaluation (#487). The single-task ACT, DP3 and RDT checkpoints were deleted from the Hugging Face repo on 2026-09-03. Users reported large gaps when reproducing π0 (for example Lift Pot 84% on the board vs 25%, #232, closed as solved) and DP3 (#441, open). Maintainers reproduced four tasks within 4 points of the board in 2025-10 (#215).
Bundled 3D models have non-commercial terms despite an MIT label
The data is labelled MIT, but some of the bundled 3D models come with non-commercial terms. Inferred13132+1
Details
The data repository is labelled MIT but includes PartNet-Mobility models, whose terms limit use to non-commercial research and education and require users to accept them, and Objaverse objects with per-object licences, some non-commercial.
DP3 uses information a real robot would not have
DP3's Easy score relies partly on perfect depth data and perfect separation of objects from the background, which only the simulator provides.1
Details
The paper says DP3's strong clean-scene results 'partly stem from perfect point clouds and clean background segmentation in simulation'; its training setup uses precise segmentation of background and tabletop.
Details
About
- What it is
- Benchmark134
More
The paper presents a benchmark with a fixed protocol; the project runs an official leaderboard.
- Built by
- Shanghai Jiao Tong University and HKU MMLab, with 26 authors351
More
26 authors. Equally leading organisations: SJTU (MoE Key Lab of AI, AI Institute) and HKU MMLab. Other affiliations include Lumina EAI, Shanghai AI Lab, Shenzhen University, NJU, TeleAI, Fudan, SYSU, USTC, CSU, SUSTech, Tsinghua, D-Robotics, NEU and HKU-SH ICRC.
Last authors Ping Luo and Yao Mu. The licence and leaderboard contact is Tianxing Chen.
- Version
- 2.0. There are no release tags.35148
More
No release tags. Paper arXiv v1 2025-06-22, v2 2025-08-27; ICML 2026. Evaluation moved to the XPolicyLab interface on 2026-08-03.
RoboTwin 1.0 and early version · Earlier benchmarks by the same group: 'RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins', CVPR 2025 (arXiv 2504.13059), and an early version at an ECCV 2024 workshop (arXiv 2409.02920). Kept on separate branches.1436
Leaderboard changes · Launched 2025-08-06 with single-task baselines; co-train and single-task results merged into one board 2026-07-22; default ranking set to the mean of Easy and Hard 2026-08-27.142
Data and reproducibility updates · Official demo_clean training data refreshed 2026-09-14 (standard RGB images, 'legal-joint replacements'); seed reproducibility fixes merged 2026-09-24.1517
Setup
- Runs in
- Simulation1
More
Real-robot runs in the paper test whether its synthetic data helps real policies; they are not a benchmark track.
- Simulator
- SAPIEN 3.0391
More
SAPIEN 3.0.0b1 physics and rendering; mplib 0.2.1 and cuRobo motion planning
- Robot
- Two arms12
More
Data and tasks support five dual-arm platforms; the benchmark and leaderboard score only Aloha-AgileX.
- Robot model
- Scores use only the Aloha-AgileX robot1231
More
Benchmark robot: Aloha-AgileX dual arm (simulated). The generator supports Aloha-AgileX, Piper, Franka, UR5 and ARX-X5 pairs.
Per-task data archives cover 3 to 5 robots (for example, handover_block has no Franka or UR5 data; adjust_bottle has no Piper data), matching the paper's lower expert success on some robots.
- Setting
- Tabletop Inferred1
More
Read from the task list and randomization settings.
- Training data
- Over 100,000 scripted demonstrations131
More
Over 100,000 scripted expert trajectories over 50 tasks. Per task and robot: 50 clean and 500 randomized.
Demonstrations come from task programs written with an LLM and checked in simulation, not from human teleoperation. 11,000 background textures generated with Stable Diffusion v2 and filtered by people.
- Changes at test
- Random clutter, lighting and backgrounds, and unseen instructions, in the Hard setting401
More
Hard setting: random backgrounds, clutter, lighting, table height (up to 3 cm) and unseen instructions. Objects start in new places in every episode.
Official configs: demo_clean (Easy) evaluates with seen instructions and no randomization; demo_randomized (Hard) uses unseen instructions, random backgrounds (2% clean), cluttered table, table height up to 0.03 m and random light. Under the official protocol policies train on clean data only, so everything randomized is unseen in training.
Scoring and access
- Scored by
- Success rate12
More
Success is a task-specific predicate checked by the simulator; an episode counts as a success the first time the predicate holds.
- Score
- Success rate in the Easy and Hard settings23441
More
Success rate per task over 100 trials, in Easy (clean) and Hard (randomized). The leaderboard ranks by the mean of the Easy and Hard averages.
Official protocol · Train on 50 clean demonstrations per task (2,500 in total) on Aloha-AgileX. Co-train track: one policy for all 50 tasks; Single track: one checkpoint per task.234
Second protocol (50 clean + 500 randomized) · Introduced by Motus (2025-12): one policy trained on 2,500 clean plus 25,000 randomized demonstrations. Used by Fast-WAM, MotuBrain and others and by the 2026 audit. Not on the official board.31112+1
Test episodes the expert can solve · For each candidate seed (starting at 100000), the scripted expert runs first; seeds where it fails or the scene is unstable are skipped. Only expert-solvable scenes are scored.41
- Who runs it
- Both teams and organisers234
More
Of 24 board entries, 16 list the RoboTwin Team as the team that produced the numbers and 8 list outside teams. The page says results are reproduced via the XPolicyLab interface and requires public code, weights and a technical report. Numbers in papers are self-reported.
- Error bars
- Not reported Inferred21
More
The paper and the leaderboard give single success rates without error bars or seed spread.
- Code licence
- MIT6
More
LICENSE: MIT, 'Copyright (c) 2025 Tianxing Chen'.
- Data licence
- MIT, as labelled on the dataset card38
More
MIT on the Hugging Face dataset card (metadata only; the card has no other text)
Apache-2.0 (third-party copy) · LeRobot's re-hosted copy lerobot/robotwin_unified is described as Apache 2.0 in LeRobot's docs.25
- Asset licence
- The MIT-labelled data repo bundles 44 PartNet-Mobility objects, whose source terms allow non-commercial research and education only, and 153 Objaverse objects, which carry per-object Creative Commons licences, some non-commercial. No licence is stated for the 534 in-house Rodin models. Inferred13132+1
More
Objaverse card: ODC-By 1.0 for the collection; objects under CC-BY (721K), CC-BY-NC (25K), CC-BY-NC-SA (52K), CC-BY-SA (16K) or CC0 (3.5K). Which licences RoboTwin's 153 objects carry was not checked. Not legal advice.
- Access
- Open to download3837
More
Code on GitHub; data, assets and co-train checkpoints in an ungated Hugging Face repo.
- Commercial use
- Unclear Inferred63832+1
More
Code and data labels (MIT) allow commercial use, but bundled third-party 3D assets come with non-commercial source terms (facts.license_assets). Not legal advice.
- Published at
- Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13673-136998
Sources 41
- 1RoboTwin 2.0 full text v2Paper · Aug 2025 · checked 10 Oct 2026
- 2RoboTwin 2.0 leaderboard data file (24 entries, per-task Easy and Hard, news)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 3Motus: A Unified Latent Action World Model (RoboTwin 2.0 protocol, Table 13)Paper · Dec 2025 · checked 10 Oct 2026
- 4What Are We Actually Benchmarking in Robot Manipulation? (Table 1, Section 4, Appendix A.1)Paper · Jun 2026 · checked 10 Oct 2026
- 5Audit data file: stat_sig_robotwin2_hard_randomized_cutoff_data.csv (19 comparisons)Repository · 2026 · checked 10 Oct 2026
- 6RoboTwin LICENSE (MIT)Repository · 2025 · checked 10 Oct 2026
- 7GigaBrain-0.7 technical report (Table 9, Appendix A.1)Paper · Aug 2026 · checked 10 Oct 2026
- 8RoboTwin 2.0, ICML 2026 proceedings page (PMLR 306:13673-13699)Paper · 2026 · checked 10 Oct 2026
- 9WorldEval: World Model as Real-World Robot Policies Evaluator (Section 4.2, real-to-sim comparison on RoboTwin 1.0)Paper · May 2025 · checked 10 Oct 2026
- 10Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation (comparison table)Secondary · May 2026 · checked 10 Oct 2026
- 11Fast-WAM: Do World Action Models Need Test-time Future Imagination? (Table 1)Paper · Mar 2026 · checked 10 Oct 2026
- 12MotuBrain: An Advanced World Action Model for Robot Control (Table 3)Paper · Apr 2026 · checked 10 Oct 2026
- 13X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment VLA Model (Table 2, Table 16)Paper · Oct 2025 · checked 10 Oct 2026
- 14RoboTwin-Platform/RoboTwin README (update log, data notes)Repository · Sep 2026 · checked 10 Oct 2026
- 15TianxingChen/RoboTwin2.0 commit history (data refresh 2026-09-14, checkpoint deletions 2026-09-03)Repository · 22 Sep 2026 · checked 10 Oct 2026
- 16RoboTwin 2.0 full text v1Paper · Jun 2025 · checked 10 Oct 2026
- 17RoboTwin PR #509 'make a given seed reproduce the same episode' (merged 2026-09-24; refs issue #288, PR #490)Repository · 24 Sep 2026 · checked 10 Oct 2026
- 18RoboTwin issue #505: same seed can produce different instructions (closed)Repository · 15 Sep 2026 · checked 10 Oct 2026
- 19RoboTwin issue #522: batch evaluation can select different seed sets (open)Repository · 9 Oct 2026 · checked 10 Oct 2026
- 20RoboTwin issue #493: success may be triggered by transient states (open)Repository · 25 Aug 2026 · checked 10 Oct 2026
- 21RoboTwin issue #491: post-settle stability for block-stack success (open)Repository · 23 Aug 2026 · checked 10 Oct 2026
- 22RoboTwin issue #506: rotate_qrcode instruction vs success check (closed 2026-09-22)Repository · 15 Sep 2026 · checked 10 Oct 2026
- 23RoboTwin issue #521: hanging_mug instructions assign the wrong arm (open)Repository · 9 Oct 2026 · checked 10 Oct 2026
- 24RoboTwin PR #488: pre-set task arm tags (merged 2026-08-20) and envs/open_laptop.pyRepository · 20 Aug 2026 · checked 10 Oct 2026
- 25LeRobot documentation: RoboTwin 2.0Repository · 20 Apr 2026 · checked 10 Oct 2026
- 26RoboTwin issue #438: possible RGB/BGR inversion in HDF5 images (open)Repository · 14 Apr 2026 · checked 10 Oct 2026
- 27RoboTwin issue #487: release the π0 leaderboard checkpoints (maintainer reply)Repository · 3 Sep 2026 · checked 10 Oct 2026
- 28RoboTwin issue #232: reproducing π0 leaderboard resultsRepository · Nov 2025 · checked 10 Oct 2026
- 29RoboTwin issue #441: DP3 success rates vs the paper (open)Repository · Apr 2026 · checked 10 Oct 2026
- 30RoboTwin issue #215: data generated differently between versions (maintainer reproduction)Repository · Oct 2025 · checked 10 Oct 2026
- 31TianxingChen/RoboTwin2.0 file tree (per-task archives, assets, co-train checkpoints)Repository · Sep 2026 · checked 10 Oct 2026
- 32PartNet-Mobility terms of use (sapien-sim/PartNetMobility gate text)Repository · Jul 2026 · checked 10 Oct 2026
- 33allenai/objaverse dataset card (licence section)Repository · 2023 · checked 10 Oct 2026
- 34RoboTwin 2.0 leaderboard page (setting, listing policy)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 35RoboTwin 2.0 (arXiv abstract page, authors, submission history)Paper · Jun 2025 · checked 10 Oct 2026
- 36Semantic Scholar record for RoboTwin 1.0 (arXiv:2504.13059, CVPR 2025)Index · 10 Oct 2026 · checked 10 Oct 2026
- 37GitHub API: RoboTwin-Platform/RoboTwin (stars, forks, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 38TianxingChen/RoboTwin2.0 dataset card and Hub API (licence, gated, downloads)Repository · 22 Sep 2026 · checked 10 Oct 2026
- 39RoboTwin scripts/requirements.txt (sapien==3.0.0b1, mplib==0.2.1)Repository · 2026 · checked 10 Oct 2026
- 40RoboTwin task configs demo_clean.yml and demo_randomized.ymlRepository · 2026 · checked 10 Oct 2026
- 41RoboTwin evaluation script scripts/eval_policy_xpolicylab.py (seeds, expert check, test_num)Repository · Sep 2026 · checked 10 Oct 2026
Where we searched for missing information
validity (paired sim-versus-real evaluations on RoboTwin 2.0): RoboTwin 2.0 paper v1 and v2 (real runs test data transfer only), ICML page, README, leaderboard page and data file (its RoboDojo tabs belong to a separate benchmark), CVPR 2025 challenge report (real round used different tasks; no sim-real comparison), WorldEval (RoboTwin 1.0, custom tasks), WorldArena and BWM (world models compared with the RoboTwin simulator, not real robots), 'Toward Visually Realistic Simulation' (2605.06311; its comparison table marks RoboTwin as having no validated sim-to-real correlation), the 2026 audit (no sim-real test). One web search (extended) for RoboTwin 2.0 sim-real correlation studies; the shared web-search budget then ran out, so later papers were not searched again.
trials and protocol details: Official eval script (scripts/eval_policy_xpolicylab.py), task configs demo_clean.yml and demo_randomized.yml, leaderboard data file (per-task values checked for parity), GigaBrain-0.7 report.
license_assets: Repo LICENSE, Hugging Face card (metadata only), Hugging Face file tree (objects.zip, background_texture.zip, embodiments.zip), RoboTwin docs index, PartNet-Mobility terms (sapien-sim Hugging Face gate; sapien.ucsd.edu refused connection), Objaverse card.
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
- Expanded to a full entry from primary sources: paper v1 and v2, ICML page, repo code and configs, leaderboard data, Hugging Face data history, GitHub issues, the 2026 audit and its data file, model papers reporting results, challenge documents and asset licence terms. Added the two-protocol problem, reproducibility and success-check issues, top scores, challenges and derived benchmarks. Corrected: LeRobot's open_laptop note is outdated (fixed upstream 2026-08-20). WorldEval's r 0.411 is kept out of validity because it used RoboTwin 1.0.