ManiSkill3
ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
ManiSkill3 is an open-source robot simulator that runs many copies of a task in parallel on one GPU. Its documentation lists 51 tasks, and it has no single benchmark score.123
What a score here does not tell you Inferred
- How well a policy (the robot's control model) will do on a real robot.The authors checked this on 4 digital twins (simulated copies of real test setups) and 1 cube task.
- How methods compare on one shared score.ManiSkill3 has no headline score and no fixed test protocol.
- Whether results hold across simulator settings.Switching between CPU and GPU simulation, or changing the package version, can change scores.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| ManiSkill3 port of SIMPLER Bridge twins Oct 2024 | Pearson r = 0.9284. MMRV (a measure of how often two rankings disagree) is 0.0147.145 | Octo-Base, Octo-Small and RT-1-X were run on ManiSkill3's GPU port of four SIMPLER WidowX tasks and compared with their real success rates on the same tasks, which equal SIMPLER's published real results. The study’s authors described the result as “close to the original values reported in SIMPLER”. | The benchmark’s authors |
| Koch cube-picking sim2real May 2025 | No statistic was reported. Success curves in simulation and on the real robot are plotted in Figure 13.1 | 14 checkpoints from each of 3 RL training runs were evaluated 8 times each in simulation and on a real Koch arm on the same cube-picking task. The study’s authors described the result as “good correlation”. | The benchmark’s authors |
Benchmarks built on ManiSkill3: ManiSkill-HAB, SimplerEnv ManiSkill3 twins, RL4VLA task set, MIKASA-Robo-VLA, RoboFactory.171811+1
Our assessment Opinion
Compare ManiSkill3 numbers only within one paper or protocol.
Reasoning
There is no single ManiSkill3 score. A number from ManiSkill3 only means something next to the task, demonstration source, controller, success metric, backend and version used. Compare results only within one paper or under a shared protocol such as RDT's five tasks.
Confidence: high
The real-world checks cover four digital twins and one cube task.
Reasoning
The real-world evidence is narrow and comes from the authors. It covers a re-implementation of four SIMPLER twins (simulated copies of real test setups), checked against SIMPLER's existing real data for three policies, and one cube task on a low-cost arm. It does not show that scores on the other ManiSkill3 tasks predict real performance.
Confidence: medium
It is best used as a fast testbed during development. It does not rank general skill.
Reasoning
ManiSkill3's main value is speed. Many copies of a task run in parallel on one GPU, which suits reinforcement learning and large data generation. Its tasks are better read as a development testbed than as a ranking of general skill.
Confidence: medium
Check asset licences before commercial use.
Reasoning
Teams planning commercial work should check asset licences task by task. The code is permissive, but the README puts the assets under a non-commercial licence.
Confidence: high
Known problems 5
Asset licences are non-commercial and inconsistent
The README puts all assets under CC BY-NC 4.0. The labels on individual assets differ.32021+2
Details
The README puts all assets under CC BY-NC 4.0. Robot folders carry their own BSD or Apache-2.0 licences, and Hugging Face copies of scene datasets are labelled CC BY 4.0 or MIT; the MIT label on the RoboCasa copy differs from RoboCasa's own CC BY 4.0. A file under LGPL-3.0 sat in the Apache-2.0 code until 2025-09-14.
CPU and GPU simulation can give different results
One policy scored 70% in CPU simulation and 12% in GPU simulation.23246
Details
A user reported a Diffusion Policy at 70% success on the CPU backend and 12% on the GPU backend over 100 episodes of a custom task, and official motion planners failing on GPU for StackCube, PegInsertionSide, PlaceSphere and LiftPegUpright (version 3.0.0b21). The maintainer replied that the planners were not fully tested on GPU and suggested collecting and testing on CPU; the maintainer closed the issue on 2026-03-14 without a code fix. The docs advise evaluating on the same backend the demonstrations were collected on. An independent test (GPUSimBench, ManiSkill 3.0.0b22) found that parallel GPU environments started from identical states end in measurably different places (mean pairwise distance 4.76 cm), while whole runs repeat exactly.
Scores depend on the package version
Beta versions changed the random-seed and evaluation code. Results therefore depend on the version.252627+1
Details
Beta releases changed evaluation-relevant code: 3.0.0b11 sometimes did not randomize environments on reset (fixed in b12, 2024-10-29); b10 aligned evaluation metrics and baseline reporting; a 2026-06-23 fix corrected seeding that shrank the episode seed batch after an unseeded reset. The SimplerEnv port notes its evaluation 'is not deterministic and results may vary between runs', and the original SIMPLER results need the older main branch.
The same method gets very different numbers
Depending on the paper, Diffusion Policy scores 0.76 or 0.40 on PickCube.11014
Details
Diffusion Policy on PickCube: 0.76 (ManiSkill3 paper, RGB, 100 demos, best during training) and 1.00 with 1,000 demos, but 40.0% in RDT's five-task table (5,000 motion-planning demos, joint position control, checkpoint picked by validation loss). On StackCube the same rows read 0.61 and 80.0%. E0's reprint of RDT's table changes OpenVLA's StackCube cell from 8% to 80.0 while keeping the 4.8% average.
There is no fixed protocol or headline score
Papers choose their own tasks, data and success metric.293024+2
Details
ManiSkill3 is a task library, and papers pick their own tasks, demonstrations, controllers and episode counts. The docs' RL 'standard benchmark' says it has 'a small set of 8 tasks' but lists 7, one under an ID that is not registered (HumanoidPlaceAppleInBowl-v1; the registered task is UnitreeG1PlaceAppleInBowl-v1); the 50-task large set 'is still being developed'. The learning-from-demonstrations baselines page is marked work in progress. Each episode records two success metrics, success_once and success_at_end, and the docs say demonstration-learning work typically reports success_once.
Details
About
- What it is
- A simulator with a library of tasks Inferred133
More
The paper's conclusion calls ManiSkill3 a 'framework/benchmark'; the GitHub description says 'an open source GPU parallelized robotics simulator and benchmark'. We file it as a simulator that hosts many tasks because it has no single fixed task set or score.
- Built by
- UC San Diego, Hillbot, Carnegie Mellon University, TU Dresden, Tsinghua University, King's College London1
More
23 authors; Hao Su is last author. 19 list UC San Diego, and 7 of those also list Hillbot. Hillbot and the Qualcomm Embodied AI Fund are thanked for support.
UC San Diego · Hao Su lab; 19 of 23 authors1
Hillbot · Robotics company; second affiliation of 7 authors including Stone Tao, Fanbo Xiang and Hao Su1
Carnegie Mellon University · Chen Bao1
TU Dresden · Roberto Calandra1
Tsinghua University · Rui Chen1
King's College London · Shan Luo1
- Released
- March 2024. The paper followed in October 2024 and was published at RSS 2025.343536+3
More
First 3.x package on PyPI on 2024-03-09 (3.0.0.dev0); first GitHub release 2024-03-28; arXiv v1 2024-10-01; RSS 2025.
arXiv v2 (2025-05-30) is the RSS 2025 demonstration-track text with a new title. The paper was also an oral at the ICLR 2025 Robot Learning Workshop.
- Version
- 3.0.1, from April 2026353440
More
3.0.1, released 2026-04-21, is the first release without the 'beta' label. Betas ran from 3.0.0.b2 (2024-05-02) to 3.0.0b22 (2025-12-05).
There is no 3.0.0 release: PyPI goes from 3.0.0b22 to 3.0.1. A tag v3.0.0b23 exists without a release. The roadmap plans to replace the PhysX backend with Newton/MuJoCo Warp, which would change the physics behind task scores.
3.0.0b11 (2024-10) · Release notes for b12 say an RNG bug in b11 sometimes stopped environments from randomizing on reset, and ask users to upgrade.25
3.0.0b10 (2024-10-01) · Added the four SIMPLER Bridge real-to-sim twins and aligned evaluation metrics and baseline reporting.26
Planned backend change · Roadmap: switch the PhysX backend to Newton/MuJoCo Warp.41
Setup
- Runs in
- Simulation1
- Simulator
- SAPIEN 3 with PhysX GPU simulation13
More
Rendering uses SAPIEN's parallel rasterizer; ray tracing is available without parallelization. GPU simulation is supported on Linux with NVIDIA GPUs only (README system table).
- Robot model
- 35 robot setups in the docs Inferred91
More
Docs have 35 robot pages (our count), including Franka Panda, xArm 6/7, UR10e, WidowX 250S, WidowX AI, SO100, Koch v1.1, Fetch, Google Robot, Unitree G1, H1 and Go2, ANYmal C, Allegro, D'Claw, Inspire hands and TriFinger. The paper says 20+ robots.
Counted from docs/source/robots/*/index.md on 2026-10-10. Some pages are variants of one robot (wrist camera, simplified legs).
- Setting
- Tabletop, Whole home Inferred24317
More
RoboCasa kitchens, ReplicaCAD and AI2-THOR scenes load in ManiSkill3, but the docs say none of these large scene datasets has trainable tasks with success/fail conditions yet; RoboCasaKitchen-v1 has no success condition. So 'kitchen' is not tagged.
- Tasks
- 51 listed tasks Inferred21
More
51 task entries in 9 documentation categories (our count). The paper groups tasks into 12 categories.
Counted 2026-10-10 from the generated task tables and pages: table-top 21, control 9, humanoid 4, digital twins 4, quadruped 3, mobile manipulation 3, dexterous 3 task families (some with difficulty levels 0 to 4), drawing 3, external 1 (ManiSkill-HAB). Of the 44 entries in task tables, 35 have success/fail conditions and 17 have demonstrations.
- Scenes
- 3 scene datasets4344
More
3 room-scale scene datasets: ReplicaCAD, AI2-THOR-family scenes (for example ArchitecTHOR) and RoboCasa kitchens
Used for exploration and data generation. Only ManiSkill-HAB defines scored tasks in ReplicaCAD.
- Training data
- Demonstrations for 16 tasks Inferred45461
More
Official demonstrations for 16 tasks on Hugging Face, about 0.47 GB of compressed files. Motion-planning sets hold 1,000 trajectories per task (two tasks checked). The paper says 'millions of demonstration frames'.
Task count and size summed by us from the Hugging Face file listing (16 zip files, 473,816,098 bytes). Episode counts read from trajectory.json for PickCube-v1 and PegInsertionSide-v1. Files store states and actions only; observations are regenerated with the replay tool.
Sources · Motion planning, RL policies and teleoperation, labelled in each file's metadata (for example 10 teleoperated PickCube-v1 demos)4546
Per-backend files · File names record the controller and the simulation backend (physx_cpu or physx_cuda)2445
- Changes at test
- Start positions vary, and some tasks also vary object shapes. Inferred2301
More
Start states are randomized every episode. Some tasks also change object geometry (for example peg shapes in PegInsertionSide). No held-out object split is documented.
The docs' evaluation setup re-randomizes every episode (reconfiguration_freq=1), which 'randomizes object geometries if the task has object randomization'. We found no statement that evaluation objects are unseen in training.
Scoring and access
- Scored by
- Success rate, Reward302
More
Tasks define success/fail conditions and, for most, dense rewards for RL. There is no benchmark-wide headline score.
- Score
- Success for each task, with two success metrics3024
More
Each episode records success_once (succeeded at any step), success_at_end (succeeded at the last step), the matching fail flags and the return. The docs say learning-from-demonstration work typically reports success_once. No benchmark-wide average.
The two success metrics can differ for the same rollouts, so a paper must say which one it reports.
- Trials
- Not fixed. It varies by paper.30241+1
More
The docs give an evaluation script (no early resets, re-randomize every episode, 64 parallel environments in the example) but no required number of episodes.
ManiSkill3 paper, imitation baselines · Best success rate obtained during training (Tables II and III); PerAct evaluated over 100 episodes1
RDT-1B five-task protocol · 250 trials per task (10 seeds x 25), mean and standard deviation10
RL baselines · Evaluation curves shared as Weights & Biases reports; RL standard benchmark small set listed in the docs29
- Who runs it
- Each team tests its own model Inferred329
More
No organiser runs submissions. Each paper runs its own evaluation.
- Error bars
- Sometimes reported Inferred11014
More
The paper's RL curves show 95% confidence intervals over 5 seeds and the Koch sim2real curves show 95% intervals; its imitation tables give single numbers. RDT reports mean and standard deviation over 10 seeds; E0 gives single numbers.
- Leaderboard
- None. Scores are only in papers. Inferred3292
More
No leaderboard in the README, docs or paper. The authors share baseline training runs as Weights & Biases reports, which are not a results table for outside methods.
- Code licence
- Apache-2.015316
More
LICENSE is Apache-2.0. Since 2025-09-14 a LICENSE-3RD-PARTY file adds the BSD licence of copied PyTorch3D code, after issue #1208 flagged a file under LGPL-3.0. The README says rigid-body environments use 'fully permissive licenses (e.g., Apache-2.0)'.
- Data licence
- Demonstration dataset card on Hugging Face: apache-2.045
More
The card covers haosulab/ManiSkill_Demonstrations. The README's asset licence (CC BY-NC 4.0) may also apply to rendered observations of those assets; the card does not say.
- Asset licence
- CC BY-NC 4.0 according to the README. Other labels conflict with it.32021+1
More
README: 'The assets are licensed under CC BY-NC 4.0'. Individual sources carry other labels (items).
CONFLICT between the blanket README statement and per-source labels. Terms for PartNet-Mobility assets (downloaded from UCSD storage) could not be checked: the SAPIEN site did not respond on 2026-10-10.
Robot models · Own LICENSE files: Allegro (BSD-style, SimLab), D'Claw (Apache-2.0, ROBEL), Unitree G1 (BSD-3-Clause), Koch (Apache-2.0, Hugging Face), SO100 (Apache-2.0)20
Scene copies on Hugging Face · haosulab/ReplicaCADRearrange and haosulab/AI2THOR cards say cc-by-4.0; haosulab/RoboCasa says mit, while RoboCasa's own README says CC BY 4.021
- Access
- Open. The package installs with pip, and the data is on Hugging Face.33445+2
More
pip install mani_skill; assets and demonstrations download from public Hugging Face and UCSD links without registration. Since a 2026-06-07 link update, the docs point to mani-skill/ManiSkill_Demonstrations, which returns HTTP 401; the download tool still uses the working haosulab repository.
- Commercial use
- Not allowed Inferred31520
More
Code (Apache-2.0) allows commercial use. The README licenses the bundled assets under CC BY-NC 4.0, which bars commercial use, even though some individual assets carry permissive labels. A company could use the simulator with its own assets. Not legal advice.
Sources 48
- 1ManiSkill3 paper, full text v2 (retitled 'Demonstrating GPU Parallelized Robot Simulation and Rendering for Generalizable Embodied AI with ManiSkill3')Paper · May 2025 · checked 10 Oct 2026
- 2ManiSkill documentation: Tasks (index and 9 category pages; generated task tables)Official site · 2026 · checked 10 Oct 2026
- 3ManiSkill README (licence section, RSS 2025, ManiSkill2 pointer)Repository · Aug 2026 · checked 10 Oct 2026
- 4ManiSkill3 paper v2, HTML version, Figure 25 image (real vs sim success for octo_base, octo_small, rt-1x)Paper · May 2025 · checked 10 Oct 2026
- 5Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER), Table V and Figure 7Paper · May 2024 · checked 10 Oct 2026
- 6GPUSimBench: Towards Scalable and Reliable GPU-Accelerated Simulators in Embodied AI (Table IV)Paper · Jul 2026 · checked 10 Oct 2026
- 7Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators (SureSim)Paper · Oct 2025 · checked 10 Oct 2026
- 8Squint: Fast Visual Reinforcement Learning for Sim-to-Real RoboticsPaper · Feb 2026 · checked 10 Oct 2026
- 9ManiSkill documentation: Robots (35 robot pages)Official site · 2026 · checked 10 Oct 2026
- 10RDT-1B README, 'Simulation Benchmark' section (ManiSkill five-task results)Repository · Dec 2024 · checked 10 Oct 2026
- 11What Can RL Bring to VLA Generalization? An Empirical Study (RL4VLA)Paper · May 2025 · checked 10 Oct 2026
- 12RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action ModelsPaper · Oct 2025 · checked 10 Oct 2026
- 13piRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (ManiSkill benchmark, Table 4)Paper · Oct 2025 · checked 10 Oct 2026
- 14E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion (Table 10)Paper · Nov 2025 · checked 10 Oct 2026
- 15ManiSkill LICENSE (Apache-2.0)Repository · 2024 · checked 10 Oct 2026
- 16ManiSkill pull request #1269 '[Docs] Licensing' (adds LICENSE-3RD-PARTY, rewrites an LGPL-derived file)Repository · 14 Sep 2025 · checked 10 Oct 2026
- 17ManiSkill-HAB: A Benchmark for Low-Level Manipulation in Home Rearrangement TasksPaper · Dec 2024 · checked 10 Oct 2026
- 18ManiSkill documentation: Digital Twins (BridgeData v2 evaluation twins)Official site · 2026 · checked 10 Oct 2026
- 19ManiSkill documentation: Community ProjectsOfficial site · Jun 2026 · checked 10 Oct 2026
- 20ManiSkill robot asset folders with their own LICENSE files (allegro, dclaw, g1_humanoid, koch, so100)Repository · 2026 · checked 10 Oct 2026
- 21Hugging Face dataset cards for ManiSkill scene copies (haosulab/ReplicaCADRearrange, haosulab/AI2THOR, haosulab/RoboCasa)Repository · 2024 · checked 10 Oct 2026
- 22ManiSkill issue #1208 'Licensing mismatch'Repository · 30 Jul 2025 · checked 10 Oct 2026
- 23ManiSkill issue #1380 'Significant discrepancy between CPU and GPU simulation backends'Repository · 25 Jan 2026 · checked 10 Oct 2026
- 24ManiSkill documentation: Learning from Demonstrations Setup (evaluation, success_once, backend pitfalls)Official site · 2026 · checked 10 Oct 2026
- 25ManiSkill release v3.0.0b12 notes (RNG bug introduced in 3.0.0b11)Repository · 29 Oct 2024 · checked 10 Oct 2026
- 26ManiSkill release v3.0.0b10 notes (real2sim twins added; evaluation metrics aligned)Repository · 1 Oct 2024 · checked 10 Oct 2026
- 27ManiSkill pull request #1457 'Fix scalar-seeded main RNG expansion shrinking the episode seed batch'Repository · 23 Jun 2026 · checked 10 Oct 2026
- 28SimplerEnv README, maniskill3 branchRepository · 2024 · checked 10 Oct 2026
- 29ManiSkill documentation: Reinforcement Learning Baselines ('Standard Benchmark')Official site · 2026 · checked 10 Oct 2026
- 30ManiSkill documentation: Reinforcement Learning Setup (evaluation protocol and metrics)Official site · 2026 · checked 10 Oct 2026
- 31ManiSkill documentation: Learning from Demonstrations Baselines (marked WIP)Official site · 2026 · checked 10 Oct 2026
- 32ManiSkill source: humanoid_pick_place.py (registers UnitreeG1PlaceAppleInBowl-v1; HumanoidPlaceAppleInBowl is a class name)Repository · 2026 · checked 10 Oct 2026
- 33GitHub API: mani-skill/ManiSkill (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 34PyPI: mani-skill release history (3.0.0.dev0 on 2024-03-09 to 3.0.1)Repository · 21 Apr 2026 · checked 10 Oct 2026
- 35ManiSkill GitHub releases and tags (v0.2.0 to v3.0.1)Repository · 21 Apr 2026 · checked 10 Oct 2026
- 36ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (arXiv abstract page, v1 and v2 history)Paper · Oct 2024 · checked 10 Oct 2026
- 37ManiSkill3 paper, full text v1Paper · Oct 2024 · checked 10 Oct 2026
- 38Robotics: Science and Systems XXI, paper 21 (proceedings page)Paper · Jun 2025 · checked 10 Oct 2026
- 39ICLR 2025, 7th Robot Learning Workshop, oral listing for ManiSkill3Paper · Apr 2025 · checked 10 Oct 2026
- 40ManiSkill release v3.0.1 notesRepository · 21 Apr 2026 · checked 10 Oct 2026
- 41ManiSkill documentation: RoadmapOfficial site · Jun 2026 · checked 10 Oct 2026
- 42ManiSkill commit history, main branchRepository · 2 Aug 2026 · checked 10 Oct 2026
- 43ManiSkill documentation: Scene DatasetsOfficial site · 2026 · checked 10 Oct 2026
- 44ManiSkill asset registry source (mani_skill/utils/assets/data.py)Repository · 2026 · checked 10 Oct 2026
- 45Hugging Face dataset haosulab/ManiSkill_Demonstrations: card and file listingRepository · 6 Jul 2025 · checked 10 Oct 2026
- 46ManiSkill demonstration metadata: PickCube-v1 motion-planning trajectory.json (1,000 episodes)Repository · 2025 · checked 10 Oct 2026
- 47ManiSkill commit bbf04bd4 'Update links to new github org' (changes the docs' Hugging Face demo link)Repository · 7 Jun 2026 · checked 10 Oct 2026
- 48Hugging Face dataset mani-skill/ManiSkill_Demonstrations (link in the docs; HTTP 401)Repository · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: ManiSkill3 paper v1 and v2 full text (Sections III-E, IV-D, Appendix VII-K, VIII, Figures 13 and 25); docs digital-twin and sim2real pages; SimplerEnv maniskill3 README; SIMPLER paper Table V; SureSim (2510.04354; custom twin); Squint (2602.21203; custom tasks); GPUSimBench (2607.13059; physics only); 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366; cites ManiSkill in related work only); PolaRiS (2512.16881), Betting for Sim-to-Real (2604.24018), Robot Policy Evaluation for Sim-to-Real Transfer (2508.11117), Active Real-World Factor-Based Evaluation (2607.14439), Beyond Binary Success (2603.13616): no ManiSkill3 mention in full text; the 2026 audit (2606.04233) does not cover ManiSkill3; web searches for ManiSkill3 sim-vs-real correlation and critiques.
objects: Paper, README, docs task pages and asset registry: no object count.
license_assets (PartNet-Mobility terms): sapien.ucsd.edu did not respond on 2026-10-10.
used_by (count): No tracker of papers reporting ManiSkill3 results found; the 2026 audit tracks five other benchmarks only.
leaderboard: README, docs (tasks, baselines, community projects), paper.
Change history
- Created full entry from primary sources, starting from the checked basic entry and the core-sim-a inventory notes. Corrected: robots (35 configurations in docs), scene (kitchen scenes have no scored tasks), leaderboard (paper-only), commercial use (non-commercial per README asset licence), real-to-sim study details (3 policies, real values taken from SIMPLER), latest commit 2026-08-02. Added lineage to ManiSkill 1 and 2, validity items, issues and readings.