RoboTwin 2.0

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

How to read this picture

RoboTwin 2.0 is a simulated benchmark and data generator for two-armed robots. It tests robot policies (the models that control a robot) on 50 tabletop tasks in clean and randomized scenes, and it has an official leaderboard.123

Sources
Last checked 10 Oct 2026Full entry59 of 69 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Not checked
Skill
Handling objects
Robot
Two arms12
Used by
Widely reported. 24 models are on the official leaderboard.425
595 citations
Licence
MIT6
Commercial use: unclear

What a score here does not tell you Inferred

  1. How well a policy will do on a real robot.No study has scored the same policies on RoboTwin 2.0 and on real robots.
  2. How a score compares with scores in other papers.Papers use two different training protocols, and they give very different Hard scores.
  3. How well a policy works on other robots.All scores use one simulated robot, the Aloha-AgileX.
  4. Whether the task stays done after a success is recorded.A success is recorded the first moment the goal is met, even if an object is still moving or falling.
ChartPublished scores over time
95% AND ABOVE1020304050607080901002026RDT · 24.11% · 2025-08DP3 · 30.1% · 2025-08π0 · 31.38% · 2025-08X-VLA · 44.45% · 2026-08π0.5 · 58.35% · 2026-08GigaBrain-0.7 · 67.35% · 2026-08UniWAM · 71.34% · 2026-09ME-Dex-1.0 · 78.85% · 2026-09PatchWAM · 79.14% · 2026-10RDT 24.11%PatchWAM 79.14%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.217

Comparisons with real robots

Tried on real robots, not compared Inferred189+1

Details

Data-transfer evidence only. The paper trained RDT on a real COBOT-Magic robot for 4 tasks: with 10 real demonstrations plus 1,000 RoboTwin trajectories, average success on unseen cluttered backgrounds rose from 9.0% to 42.0%; with synthetic data only it reached 29.5% (arXiv: 367% and 228% relative gains; ICML abstract: '3.6x' and '2.2x'). Trial counts for these real runs are not stated. Independent check on the earlier version only: WorldEval (2025-05) ran real-trained policies in RoboTwin 1.0 on 3 custom tasks, with images translated by MidJourney, and found average Pearson r 0.411 and MMRV 0.261 against real results. That study predates 2.0 and uses other tasks, so it is not counted as validity for 2.0. A May 2026 paper by other authors also marks RoboTwin as having no validated sim-to-real correlation in its comparison table.

Our assessment Opinion

RoboTwin 2.0 is a useful simulation test. Its link to real-robot results has not been measured.

Reasoning

RoboTwin 2.0 tests two-arm skills under visual clutter, and the 2026 audit found no cheap shortcut on it. A high score still says little about real robots. No study has scored the same policies on version 2.0 and on real hardware.

Confidence: medium

Check which training protocol a score used.

Reasoning

Before comparing numbers, check the training protocol. The official leaderboard uses 50 clean demonstrations per task. The other protocol adds 500 randomized demonstrations per task. Hard scores under the two are not comparable.

Confidence: high

Look at the Hard scores as well as the average.

Reasoning

Read Easy and Hard separately. Several policies do well in clean scenes and score near zero in randomized ones. On the leaderboard, FastWAM scores 77.8 on Easy and 1.9 on Hard. The leaderboard's average hides this.

Confidence: high

Ignore small gaps between leaderboard entries.

Reasoning

Small gaps between leaderboard entries are hard to interpret. There are no error bars, and entries were produced by different teams on different code and data versions. Some entries appear to use 50 rather than 100 trials.

Confidence: medium

Known problems 9

  1. Two training protocols give very different Hard scores

    Depending on the training data, Hard scores for the same model range from about 2% to 92%.2311+2

    Details

    The official board trains on 50 clean demonstrations per task; its best Hard score is 68.12. Since Motus (2025-12) many papers train one policy on 50 clean plus 500 randomized demonstrations per task, which puts randomized scenes into training; MotuBrain reports 96.1 Hard. The same model names differ widely: Fast-WAM 91.78 Hard in its paper vs FastWAM 1.9 Hard on the board; the audit's data file lists starVLA-α at 88.3 and ABot-M0 at 85.08, while the board shows starVLA 3.16 and Abot-M0 30.36. Both protocols are called 'RoboTwin 2.0' in papers.

  2. The same baseline gets different scores in different sources

    Different sources report the Hard score of π0.5 as 46.0, 43.84 and 76.76.2311+2

    Details

    π0.5: 70.7 / 46.0 on the official board; 42.98 / 43.84 in Motus's run and 82.74 / 76.76 in Fast-WAM's table, both under the 50+500 protocol. X-VLA: 70.0 / 39.0 in its own paper (training data not stated); 68.0 / 20.9 on the board; 72.8 / 72.84 in Motus's run. The audit replaced X-VLA's own number with a protocol-matched one because X-VLA did not train on 50 clean plus 500 randomized demonstrations.

  3. Code and data changed while the leaderboard was running

    Leaderboard entries were produced with different versions of the code and data.141516+2

    Details

    DP3 evaluation code was fixed on 2025-07-19 and ACT deployment code on 2025-08-25, with the board updated. Several success checks changed in July and August 2025. The official demo_clean training data was refreshed on 2026-09-14 with 'legal-joint replacements', after most co-train entries were listed. Paper v1 reported clean-to-randomized drops of 28.2 (RDT) and 40.3 (π0) points on 13 tasks; v2 reports 20.8 and 30.1 on 50 tasks. Board entries date from 2025-08 to 2026-10 and so span these changes.

  4. Seeds did not reproduce the same scene until September 2026

    The same seed (the number that sets up a random scene) could produce a different scene. This was fixed in September 2026. A separate bug in batch evaluation is still open.171819

    Details

    Users found that a dataset seed did not recreate the recorded scene (issue #288). Causes: asset order depended on the file system in 5 tasks, table-clutter bounds grew with each setup call, and Python's random generator, which picks the instruction, was not seeded. Fixes were merged on 2026-09-24 (PRs #490 and #509). A separate open bug (#522, 2026-10-09) shows batch evaluation can score a different set of seeds depending on the number of workers.

  5. Success checks can count brief moments and some disagree with the instructions

    A success can be recorded while an object is still falling. Some instructions ask for something different from what the success check tests.202122+3

    Details

    An episode ends as a success the first time the task predicate holds, without checking that objects stay put; users report bottles and stacked blocks counted as successes while still moving or falling (issues #493 and #491, open). rotate_qrcode instructions asked to lift and rotate while the check required putting the object down (fixed 2026-09-22); some hanging_mug instructions name the wrong arm (#521, open). LeRobot's docs say open_laptop's check fails during normal policy evaluation; RoboTwin fixed the arm-tag setup for open_laptop, place_object_scale and put_object_cabinet on 2026-08-20 (PR #488), after LeRobot's note.

  6. Image colour channels differ between data versions

    Older data stores images with the red and blue channels swapped. The official decoder handles both formats.142615

    Details

    The README says older ('legacy') data stores red and blue swapped and was never migrated, so old and new formats can appear in one training run; images must be decoded with decode_image_bit. A user reported the swap in 2026-04 (issue #438); a maintainer replied that images follow RGB order. The demo_clean data and LeRobot archives were refreshed with standard RGB images on 2026-09-14.

  7. Some leaderboard results cannot be re-checked

    The trained models (checkpoints) behind several leaderboard baselines are not available. Users report gaps when they try to reproduce some results.271528+2

    Details

    The π0 checkpoints behind the board were never released; a maintainer said π0 was not part of the latest re-evaluation (#487). The single-task ACT, DP3 and RDT checkpoints were deleted from the Hugging Face repo on 2026-09-03. Users reported large gaps when reproducing π0 (for example Lift Pot 84% on the board vs 25%, #232, closed as solved) and DP3 (#441, open). Maintainers reproduced four tasks within 4 points of the board in 2025-10 (#215).

  8. Bundled 3D models have non-commercial terms despite an MIT label

    The data is labelled MIT, but some of the bundled 3D models come with non-commercial terms. Inferred13132+1

    Details

    The data repository is labelled MIT but includes PartNet-Mobility models, whose terms limit use to non-commercial research and education and require users to accept them, and Objaverse objects with per-object licences, some non-commercial.

  9. DP3 uses information a real robot would not have

    DP3's Easy score relies partly on perfect depth data and perfect separation of objects from the background, which only the simulator provides.1

    Details

    The paper says DP3's strong clean-scene results 'partly stem from perfect point clouds and clean background segmentation in simulation'; its training setup uses precise segmentation of background and tabletop.

Details

About

What it is
Benchmark134
More

The paper presents a benchmark with a fixed protocol; the project runs an official leaderboard.

Built by
Shanghai Jiao Tong University and HKU MMLab, with 26 authors351
More

26 authors. Equally leading organisations: SJTU (MoE Key Lab of AI, AI Institute) and HKU MMLab. Other affiliations include Lumina EAI, Shanghai AI Lab, Shenzhen University, NJU, TeleAI, Fudan, SYSU, USTC, CSU, SUSTech, Tsinghua, D-Robotics, NEU and HKU-SH ICRC.

Last authors Ping Luo and Yao Mu. The licence and leaderboard contact is Tianxing Chen.

Released
June 20253514
More

Released 2025-06-21 (README); arXiv v1 2025-06-22

Version
2.0. There are no release tags.35148
More

No release tags. Paper arXiv v1 2025-06-22, v2 2025-08-27; ICML 2026. Evaluation moved to the XPolicyLab interface on 2026-08-03.

RoboTwin 1.0 and early version · Earlier benchmarks by the same group: 'RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins', CVPR 2025 (arXiv 2504.13059), and an early version at an ECCV 2024 workshop (arXiv 2409.02920). Kept on separate branches.1436

Leaderboard changes · Launched 2025-08-06 with single-task baselines; co-train and single-task results merged into one board 2026-07-22; default ranking set to the mean of Easy and Hard 2026-08-27.142

Data and reproducibility updates · Official demo_clean training data refreshed 2026-09-14 (standard RGB images, 'legal-joint replacements'); seed reproducibility fixes merged 2026-09-24.1517

Last update
October 202623738
More

Leaderboard updated 2026-10-10; last code push 2026-09-24; Hugging Face data changed 2026-09-22

Status
Active Inferred237
More

Leaderboard updated 2026-10-10; code and data changed in 2026-09; 89 open issues and pull requests.

Setup

Runs in
Simulation1
More

Real-robot runs in the paper test whether its synthetic data helps real policies; they are not a benchmark track.

Simulator
SAPIEN 3.0391
More

SAPIEN 3.0.0b1 physics and rendering; mplib 0.2.1 and cuRobo motion planning

Robot
Two arms12
More

Data and tasks support five dual-arm platforms; the benchmark and leaderboard score only Aloha-AgileX.

Robot model
Scores use only the Aloha-AgileX robot1231
More

Benchmark robot: Aloha-AgileX dual arm (simulated). The generator supports Aloha-AgileX, Piper, Franka, UR5 and ARX-X5 pairs.

Per-task data archives cover 3 to 5 robots (for example, handover_block has no Franka or UR5 data; adjust_bottle has no Piper data), matching the paper's lower expert success on some robots.

Setting
Tabletop Inferred1
More

Read from the task list and randomization settings.

Tasks
50 tasks352
Training data
Over 100,000 scripted demonstrations131
More

Over 100,000 scripted expert trajectories over 50 tasks. Per task and robot: 50 clean and 500 randomized.

Demonstrations come from task programs written with an LLM and checked in simulation, not from human teleoperation. 11,000 background textures generated with Stable Diffusion v2 and filtered by people.

Changes at test
Random clutter, lighting and backgrounds, and unseen instructions, in the Hard setting401
More

Hard setting: random backgrounds, clutter, lighting, table height (up to 3 cm) and unseen instructions. Objects start in new places in every episode.

Official configs: demo_clean (Easy) evaluates with seen instructions and no randomization; demo_randomized (Hard) uses unseen instructions, random backgrounds (2% clean), cluttered table, table height up to 0.03 m and random light. Under the official protocol policies train on clean data only, so everything randomized is unseen in training.

Scoring and access

Scored by
Success rate12
More

Success is a task-specific predicate checked by the simulator; an episode counts as a success the first time the predicate holds.

Score
Success rate in the Easy and Hard settings23441
More

Success rate per task over 100 trials, in Easy (clean) and Hard (randomized). The leaderboard ranks by the mean of the Easy and Hard averages.

Official protocol · Train on 50 clean demonstrations per task (2,500 in total) on Aloha-AgileX. Co-train track: one policy for all 50 tasks; Single track: one checkpoint per task.234

Second protocol (50 clean + 500 randomized) · Introduced by Motus (2025-12): one policy trained on 2,500 clean plus 25,000 randomized demonstrations. Used by Fast-WAM, MotuBrain and others and by the 2026 audit. Not on the official board.31112+1

Test episodes the expert can solve · For each candidate seed (starting at 100000), the scripted expert runs first; seeds where it fails or the scene is unstable are skipped. Only expert-solvable scenes are scored.41

Trials
100 per task in each setting12
More

Two board entries look like 50-trial runs · All 100 per-task values of GigaBrain-0.7 and of Discrete Forcing are even numbers; every other entry has odd values. This fits 50 trials per task, not 100.27

Who runs it
Both teams and organisers234
More

Of 24 board entries, 16 list the RoboTwin Team as the team that produced the numbers and 8 list outside teams. The page says results are reproduced via the XPolicyLab interface and requires public code, weights and a technical report. Numbers in papers are self-reported.

Error bars
Not reported Inferred21
More

The paper and the leaderboard give single success rates without error bars or seed spread.

Leaderboard
Official leaderboard234
Code licence
MIT6
More

LICENSE: MIT, 'Copyright (c) 2025 Tianxing Chen'.

Data licence
MIT, as labelled on the dataset card38
More

MIT on the Hugging Face dataset card (metadata only; the card has no other text)

Apache-2.0 (third-party copy) · LeRobot's re-hosted copy lerobot/robotwin_unified is described as Apache 2.0 in LeRobot's docs.25

Asset licence
The MIT-labelled data repo bundles 44 PartNet-Mobility objects, whose source terms allow non-commercial research and education only, and 153 Objaverse objects, which carry per-object Creative Commons licences, some non-commercial. No licence is stated for the 534 in-house Rodin models. Inferred13132+1
More

Objaverse card: ODC-By 1.0 for the collection; objects under CC-BY (721K), CC-BY-NC (25K), CC-BY-NC-SA (52K), CC-BY-SA (16K) or CC0 (3.5K). Which licences RoboTwin's 153 objects carry was not checked. Not legal advice.

Access
Open to download3837
More

Code on GitHub; data, assets and co-train checkpoints in an ungated Hugging Face repo.

Commercial use
Unclear Inferred63832+1
More

Code and data labels (MIT) allow commercial use, but bundled third-party 3D assets come with non-commercial source terms (facts.license_assets). Not legal advice.

Published at
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13673-136998

Sources 41

  1. 1RoboTwin 2.0 full text v2Paper · Aug 2025 · checked 10 Oct 2026
  2. 2RoboTwin 2.0 leaderboard data file (24 entries, per-task Easy and Hard, news)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
  3. 3Motus: A Unified Latent Action World Model (RoboTwin 2.0 protocol, Table 13)Paper · Dec 2025 · checked 10 Oct 2026
  4. 4What Are We Actually Benchmarking in Robot Manipulation? (Table 1, Section 4, Appendix A.1)Paper · Jun 2026 · checked 10 Oct 2026
  5. 5Audit data file: stat_sig_robotwin2_hard_randomized_cutoff_data.csv (19 comparisons)Repository · 2026 · checked 10 Oct 2026
  6. 6RoboTwin LICENSE (MIT)Repository · 2025 · checked 10 Oct 2026
  7. 7GigaBrain-0.7 technical report (Table 9, Appendix A.1)Paper · Aug 2026 · checked 10 Oct 2026
  8. 8RoboTwin 2.0, ICML 2026 proceedings page (PMLR 306:13673-13699)Paper · 2026 · checked 10 Oct 2026
  9. 9WorldEval: World Model as Real-World Robot Policies Evaluator (Section 4.2, real-to-sim comparison on RoboTwin 1.0)Paper · May 2025 · checked 10 Oct 2026
  10. 10Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation (comparison table)Secondary · May 2026 · checked 10 Oct 2026
  11. 11Fast-WAM: Do World Action Models Need Test-time Future Imagination? (Table 1)Paper · Mar 2026 · checked 10 Oct 2026
  12. 12MotuBrain: An Advanced World Action Model for Robot Control (Table 3)Paper · Apr 2026 · checked 10 Oct 2026
  13. 13X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment VLA Model (Table 2, Table 16)Paper · Oct 2025 · checked 10 Oct 2026
  14. 14RoboTwin-Platform/RoboTwin README (update log, data notes)Repository · Sep 2026 · checked 10 Oct 2026
  15. 15TianxingChen/RoboTwin2.0 commit history (data refresh 2026-09-14, checkpoint deletions 2026-09-03)Repository · 22 Sep 2026 · checked 10 Oct 2026
  16. 16RoboTwin 2.0 full text v1Paper · Jun 2025 · checked 10 Oct 2026
  17. 17RoboTwin PR #509 'make a given seed reproduce the same episode' (merged 2026-09-24; refs issue #288, PR #490)Repository · 24 Sep 2026 · checked 10 Oct 2026
  18. 18RoboTwin issue #505: same seed can produce different instructions (closed)Repository · 15 Sep 2026 · checked 10 Oct 2026
  19. 19RoboTwin issue #522: batch evaluation can select different seed sets (open)Repository · 9 Oct 2026 · checked 10 Oct 2026
  20. 20RoboTwin issue #493: success may be triggered by transient states (open)Repository · 25 Aug 2026 · checked 10 Oct 2026
  21. 21RoboTwin issue #491: post-settle stability for block-stack success (open)Repository · 23 Aug 2026 · checked 10 Oct 2026
  22. 22RoboTwin issue #506: rotate_qrcode instruction vs success check (closed 2026-09-22)Repository · 15 Sep 2026 · checked 10 Oct 2026
  23. 23RoboTwin issue #521: hanging_mug instructions assign the wrong arm (open)Repository · 9 Oct 2026 · checked 10 Oct 2026
  24. 24RoboTwin PR #488: pre-set task arm tags (merged 2026-08-20) and envs/open_laptop.pyRepository · 20 Aug 2026 · checked 10 Oct 2026
  25. 25LeRobot documentation: RoboTwin 2.0Repository · 20 Apr 2026 · checked 10 Oct 2026
  26. 26RoboTwin issue #438: possible RGB/BGR inversion in HDF5 images (open)Repository · 14 Apr 2026 · checked 10 Oct 2026
  27. 27RoboTwin issue #487: release the π0 leaderboard checkpoints (maintainer reply)Repository · 3 Sep 2026 · checked 10 Oct 2026
  28. 28RoboTwin issue #232: reproducing π0 leaderboard resultsRepository · Nov 2025 · checked 10 Oct 2026
  29. 29RoboTwin issue #441: DP3 success rates vs the paper (open)Repository · Apr 2026 · checked 10 Oct 2026
  30. 30RoboTwin issue #215: data generated differently between versions (maintainer reproduction)Repository · Oct 2025 · checked 10 Oct 2026
  31. 31TianxingChen/RoboTwin2.0 file tree (per-task archives, assets, co-train checkpoints)Repository · Sep 2026 · checked 10 Oct 2026
  32. 32PartNet-Mobility terms of use (sapien-sim/PartNetMobility gate text)Repository · Jul 2026 · checked 10 Oct 2026
  33. 33allenai/objaverse dataset card (licence section)Repository · 2023 · checked 10 Oct 2026
  34. 34RoboTwin 2.0 leaderboard page (setting, listing policy)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
  35. 35RoboTwin 2.0 (arXiv abstract page, authors, submission history)Paper · Jun 2025 · checked 10 Oct 2026
  36. 36Semantic Scholar record for RoboTwin 1.0 (arXiv:2504.13059, CVPR 2025)Index · 10 Oct 2026 · checked 10 Oct 2026
  37. 37GitHub API: RoboTwin-Platform/RoboTwin (stars, forks, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
  38. 38TianxingChen/RoboTwin2.0 dataset card and Hub API (licence, gated, downloads)Repository · 22 Sep 2026 · checked 10 Oct 2026
  39. 39RoboTwin scripts/requirements.txt (sapien==3.0.0b1, mplib==0.2.1)Repository · 2026 · checked 10 Oct 2026
  40. 40RoboTwin task configs demo_clean.yml and demo_randomized.ymlRepository · 2026 · checked 10 Oct 2026
  41. 41RoboTwin evaluation script scripts/eval_policy_xpolicylab.py (seeds, expert check, test_num)Repository · Sep 2026 · checked 10 Oct 2026
Where we searched for missing information

validity (paired sim-versus-real evaluations on RoboTwin 2.0): RoboTwin 2.0 paper v1 and v2 (real runs test data transfer only), ICML page, README, leaderboard page and data file (its RoboDojo tabs belong to a separate benchmark), CVPR 2025 challenge report (real round used different tasks; no sim-real comparison), WorldEval (RoboTwin 1.0, custom tasks), WorldArena and BWM (world models compared with the RoboTwin simulator, not real robots), 'Toward Visually Realistic Simulation' (2605.06311; its comparison table marks RoboTwin as having no validated sim-to-real correlation), the 2026 audit (no sim-real test). One web search (extended) for RoboTwin 2.0 sim-real correlation studies; the shared web-search budget then ran out, so later papers were not searched again.

trials and protocol details: Official eval script (scripts/eval_policy_xpolicylab.py), task configs demo_clean.yml and demo_randomized.yml, leaderboard data file (per-task values checked for parity), GigaBrain-0.7 report.

license_assets: Repo LICENSE, Hugging Face card (metadata only), Hugging Face file tree (objects.zip, background_texture.zip, embodiments.zip), RoboTwin docs index, PartNet-Mobility terms (sapien-sim Hugging Face gate; sapien.ucsd.edu refused connection), Objaverse card.

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
  2. Expanded to a full entry from primary sources: paper v1 and v2, ICML page, repo code and configs, leaderboard data, Hugging Face data history, GitHub issues, the 2026 audit and its data file, model papers reporting results, challenge documents and asset licence terms. Added the two-protocol problem, reproducibility and success-check issues, top scores, challenges and derived benchmarks. Corrected: LeRobot's open_laptop note is outdated (fixed upstream 2026-08-20). WorldEval's r 0.411 is kept out of validity because it used RoboTwin 1.0.