SimplerEnv
Evaluating Real-World Robot Manipulation Policies in Simulation
SimplerEnv is a set of simulated copies of two real robot test setups, used to score robot policies (the models that control a robot) in simulation. Its scores agreed with real-robot results for policies from 2023 and 2024, and later checks with newer policies found weaker agreement.123
What a score here does not tell you Inferred
- How well current vision-language-action (VLA) models will do on real robots.The authors validated it on policies from 2023 and 2024. Later checks with newer policies found weaker agreement.
- Whether a small improvement is real.Most claimed improvements cannot be shown to be statistically significant.
- How well a policy generalises beyond its training data.Policies trained on data close to the test, or inside the test simulator, can reach top scores.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| SIMPLER paper: Google Robot, Visual Matching May 2024 | Pearson r = 0.924 and MMRV = 0.056 (mean of 3 task groups)1 | The same 6 checkpoints (3 RT-1 checkpoints, RT-1-X, RT-2-X and Octo-Base, so 4 distinct models) were scored in simulation and on the real Google Robot, on the same tasks and trial grids. Agreement was measured with Pearson r and MMRV (a measure of how often two rankings disagree). The study’s authors described the result as “strong”. | The benchmark’s authors |
| SIMPLER paper: Google Robot, Variant Aggregation May 2024 | Pearson r = 0.778 and MMRV = 0.143 (mean of 3 task groups)1 | The same 6 checkpoints were scored in simulation with randomised backgrounds, lighting, distractors and table textures, and compared with the same real results. | The benchmark’s authors |
| SIMPLER paper: WidowX and BridgeData V2 May 2024 | Pearson r = 0.890 and MMRV = 0.014 (mean over 4 tasks, for success and grasp)1 | 3 policies (RT-1-X, Octo-Base and Octo-Small, so 2 model families) were scored in simulation and on a real WidowX on the same 4 tasks, with 24 real trials per task. The study’s authors described the result as “strong”. | The benchmark’s authors |
| SIMPLER camera-ready: Google Robot with OpenVLA added Jan 2025 | Pearson r = 0.929 and MMRV = 0.049 (Fig. 4)4 | 7 checkpoints (the 6 above plus OpenVLA-7B, so 5 distinct models) were scored in simulation (Visual Matching) and on the real Google Robot, on the same tasks. The study’s authors described the result as “strong”. | The benchmark’s authors |
| SIMPLER paper: Isaac Sim re-implementation May 2024 | Pearson r = 0.919 and MMRV = 0.058. In SAPIEN, r = 0.923 and MMRV = 0.082.1 | 5 checkpoints (3 RT-1 checkpoints, RT-1-X and Octo-Base) were scored on the pick coke can and move near tasks in Variant Aggregation scenes rebuilt in Isaac Sim, and compared with the same real results. | The benchmark’s authors |
| SIMPLER paper: sensitivity to visual shifts May 2024 | Pearson r = 0.831 and MMRV = 0.000 without augmentation. Pearson r = 0.970 and MMRV = 0.016 with augmentation.1 | 2 RT-1 policies, with and without image augmentation, were tested under 5 visual shifts. Their drop in success in simulation was compared with earlier real-world tests by Xie et al. The study’s authors described the result as “accurately reflect”. | The benchmark’s authors |
| ManiSkill3 GPU port Oct 2024 | Pearson r = 0.9284 and MMRV = 0.01475 | 3 policies (Octo-Base, Octo-Small and RT-1-X) were run on the GPU version of the 4 WidowX tasks and compared with the SIMPLER paper's real results. Success and grasp rates were pooled into one plot. The study’s authors described the result as “close to that of the original paper”. | The benchmark’s authors |
| AutoEval paper (UC Berkeley) Mar 2025 | Mean Pearson r = 0.548 and MMRV = 0.207 over 4 tasks (our computation from the paper's tables)2 | 6 policies (OpenVLA, Octo, Open-pi0, MiniVLA, SuSIE and SuSIE's low-level policy, so 5 distinct models) were run 50 times per task in SIMPLER and in human-run real evaluations on matching WidowX cells. Of the 4 tasks, only eggplant-in-basket is an official SimplerEnv task. The study’s authors described the result as “policy dependent”. | The benchmark’s authors |
| Wang et al. (Tsinghua, Shanghai Qi Zhi) Jun 2026 | Mean Spearman correlation 0.400, Pearson r = 0.402 and MMRV = 0.1283 | 5 VLA policies (pi0, pi0-FAST, pi0.5, GR00T N1.6 and GR00T N1.7) were run on 7 tabletop tasks rebuilt in SIMPLER's SAPIEN setting and matched to real tasks on DROID hardware, under changes to vision, layout and language. Each policy had 20 simulated and 5 real test runs per task and change. | An independent group |
Benchmarks built on SimplerEnv: ManiSkill3 BridgeData v2 digital twins, SimplerEnv-OpenVLA, AutoEval SIMPLER scenes, Audit stack-task variants.25262+1
Our assessment Opinion
Scores agreed with real-robot results for older policies. The evidence for current VLA models is weak.
Reasoning
A SimplerEnv score is good evidence of real-robot ranking only for policies like those its authors tested (RT-1, RT-2-X, Octo, OpenVLA) on these two setups. For current VLA models the evidence is weak. The two later studies that included newer policies found r = 0.548 and r = 0.402, both mostly on tasks outside the official set.
Confidence: medium
Check the protocol before comparing numbers across papers.
Reasoning
Do not compare WidowX numbers across papers without checking trials, step limits, training data and which baseline values were copied. Reported numbers for the same model differ by up to 42 points, and independent reruns came in lower than published.
Confidence: high
Top WidowX scores say little about real-world generalisation.
Reasoning
Above about 95% on WidowX, scores no longer separate methods, and such scores can be reached by training close to the test or inside the simulator. Scores on the Google Robot settings are further from 100%.
Confidence: medium
Its method for checking a simulator against real robots has become the common standard.
Reasoning
SimplerEnv's method for checking a simulator against real robots has become the common standard. It scores the same policies in simulation and on real robots and compares the results with Pearson r (a correlation) and MMRV (a measure of how often two rankings disagree). AutoEval, PolaRiS and later studies report the same two numbers.
Confidence: medium
Known problems 10
Most claimed improvements are not shown to be statistically significant
Only about 1 in 5 claimed improvements on WidowX can be shown to be statistically significant from the published scores.2728
Details
The 2026 audit checked 122 previous-best-to-new comparisons on the WidowX protocol, using only published scores. Its paper classes 19.7% (24) as provably significant at the 5% level, 50.8% (62) as inconclusive, and the rest as no improvement or provably not significant. On 2026-10-08 the audit authors released corrected classifications: 27 of 122 provably significant (22.1% by our arithmetic), 62 inconclusive, 16 provably not significant and 17 no improvement.
Independent reruns score lower than the published numbers
When a 2026 audit reran five published policies, they scored 2.6 to 23.4 points below the published numbers.292713+4
Details
The 2026 audit reran five published WidowX policies on the official grid with the standard step limits (each of the 24 start states repeated 12 times; 1,152 episodes per policy). Results: CogACT-Base 48.7%, SpatialVLA 37.5%, InternVLA-M1 61.2%, X-VLA 72.4% and Dexbotic DB-MemVLA 64.7%. The papers report 51.3%, 42.7% (34.4% without fine-tuning), 71.7%, 95.8% and 84.4%. On the stack task alone, X-VLA reports 95.8% and the audit measured 59.7% (172/288). The audit says several of these policies report under longer, easier episode step limits.
Training data placed next to the test can match top scores
A 2026 audit trained small models (22 million parameters) on 120 scripted demonstrations per task, recorded in simulation next to the test positions. They scored 91 of 96 trials, against 92 of 96 reported by X-VLA.27
Details
SimplerEnv tests in simulation, while its standard training data are real BridgeData V2 demonstrations, so a high score reads as transfer across that gap. The benchmark does not restrict training data. The 2026 audit trained a separate 22M-parameter policy per task, with no robotics pretraining, on 120 scripted demonstrations recorded in simulation next to the official test grid. Together the four policies scored 91/96 (94.8%), against 92/96 (95.8%) reported by X-VLA. The audit calls this an existence result: the score alone cannot tell crossing the gap from removing it. The same audit found no shortcut of the LIBERO kind: a small probe trained on the standard Bridge data scored 0.0%.
Results vary between runs and machines
One user saw the same model score between 16% and 40% on one task across runs.303127+2
Details
The lead author advises at least 75 trials per Bridge task because the 24-trial grid is too noisy. A user saw 16% to 40% success for the same OpenVLA model on one task across runs. Another user's repeated RT-1-X runs differed from the paper and from each other, although RT-1-X outputs are deterministic; the lead author attributes this to simulator nondeterminism. The 2026 audit found SimplerEnv rollouts with CogACT diverged from step 0 when only the CPU or only the GPU was changed. A 2026 issue reports the task and object positions changing between runs with a fixed seed (no reply by 2026-10-10). The GPU (ManiSkill3) version is a separate implementation; its README says reproducing the paper needs the main branch.
Small changes seen in the training data lower scores on the stack task
When the cubes started on support blocks, X-VLA's success on the stack task fell from 59.7% to 31.3%.27
Details
The 2026 audit changed the WidowX stack task in three ways that also occur in the BridgeData V2 training data: reversed colour order in the instruction, cubes starting on support blocks, and random cube and arm start poses (288 episodes per policy each). Stacked supports cut X-VLA from 172/288 to 90/288 (28.47 points, 95% CI 20.46 to 35.96) and Dexbotic from 129/288 to 79/288 (17.36 points). The reversed instruction cut CogACT by 11.11 points and InternVLA-M1 by 9.38. The audit reads these drops as overfitting to the fixed test grid.
Top WidowX scores are close to 100%
By early 2026, top WidowX averages reached 95.8% to 97.9%.172122+2
Details
Since October 2025, X-VLA (95.8%), SimVLA (95.8%) and CORAL (97.9%) report WidowX averages above 95% without RL in the simulator, and STARE-VLA reports 98.0% after RL inside SimplerEnv. Several tasks have reported scores of 100%. Google Robot results are lower: the highest we found are 85.5% (Visual Matching) and 84.9% (Variant Aggregation).
The same model gets very different scores in different papers
Different papers report pi0's WidowX average as anywhere from 27.1% to 69.2%.152019+5
Details
WidowX averages reported for pi0 by other papers: 69.2% (EO-1, Xiaomi-Robotics-0), 66.7% (FASTer), 40.1% (SimVLA), 27.8% (X-VLA) and 27.1% (InternVLA-M1). For OpenVLA: 1.0% (SpatialVLA, EO-1), 4.2% (CogACT, InternVLA-M1), 7.8% (SimVLA), 8.3% (X-VLA) and 29.5% (FASTer). Google Robot averages cover 3 task groups in some papers and 4 in others: SpatialVLA reports its own Visual Matching score as 75.1% over 3 tasks, while EO-1 lists SpatialVLA at 55.3% over 4. Some tables do not add up: InternVLA-M1 lists pi0-FAST per-task scores that average 32.1% next to an average of 48.3%; SimVLA's WidowX table gives Octo-Base per-task scores of 12.5 / 8.3 / 0.0 / 43.1 with an average of 31.3. Xiaomi-Robotics-0 (2026-02) calls its 79.2% WidowX average the best result, while its comparison table leaves out X-VLA's 95.8% from 2025-10.
Some top scores come from training inside the test simulator
Reinforcement learning (RL) inside the simulator raised pi0 from 67.2% to 86.7%.232434
Details
piRL raised pi0 from 67.2% to 86.7% on the WidowX tasks with RL inside the simulator. STARE-VLA reports 98.0% after RL fine-tuning in SimplerEnv and reports the best checkpoint under the same evaluation it publishes. World-Gymnast trained RL baselines in SIMPLER copies of the AutoEval scenes and found they transferred poorly to the real cells: for example 34% versus 58% for its world-model method on opening the drawer.
The real-robot check covered older policies on two setups
The authors' real-robot check covered policies from the RT-1 era on two setups. The policies most reported today were not part of it. Inferred1435+1
Details
The authors' validation covers RT-1 checkpoints, RT-1-X, RT-2-X, Octo and, in the camera-ready, OpenVLA on the Google Robot, and RT-1-X and Octo on WidowX. Each per-task correlation rests on 3 to 7 points. The policies most reported today (pi0, pi0.5, GR00T, X-VLA) were not part of it. The real Google Robot evaluations were run by Google staff, so outside groups cannot repeat that comparison (our inference); the WidowX setup can be rebuilt. Visual Matching pastes real images behind a fixed camera view, so it cannot serve policies that use wrist cameras, which most current generalist policies do (PolaRiS paper).
The paper's own tables and versions disagree
Different tables and versions of the SimplerEnv paper give different correlation numbers.1436
Details
In arXiv v1, Table I gives the Google Robot drawer task r 0.942 and MMRV 0.027 under Visual Matching, while Table IV and Fig. 6 give 0.915 and 0.055; the headline mean (r 0.924, MMRV 0.056) uses Table I. Fig. 8 gives MMRV 0.016 for the augmented RT-1 policy where Table VI gives 0.041. The camera-ready adds OpenVLA-7B to the Google Robot study and changes the per-task numbers (drawer r 0.823), but keeps the 6-checkpoint means in its Table 1. For put-eggplant-in-basket, Appendix B states 24 real trials, the real values imply 30, and the repo's metrics.py lists 0.250 and 0.400 for Octo-Base and Octo-Small where the paper has 0.233 and 0.433.
Details
About
- What it is
- Benchmark Inferred137
More
The paper presents a suite of simulated evaluation environments with fixed tasks, trial grids and success checks; papers report success rates on it as a benchmark. Classification by the Atlas.
- Built by
- UC San Diego, Stanford University, UC Berkeley, Google DeepMind38139
More
16 authors. Equal contribution: Xuanlin Li, Kyle Hsu, Jiayuan Gu. Co-advising: Hao Su, Quan Vuong, Ted Xiao.
UC San Diego · Hao Su's lab; lead authors and the SAPIEN/ManiSkill code base.1
Stanford University · Kyle Hsu, Chelsea Finn, Jiajun Wu and others.1
UC Berkeley · Sergey Levine's lab; WidowX real-world evaluations.1
Google DeepMind · RT-series policies; the real Google Robot evaluations were run by Google staff, per the lead author in issue #77.135
- Released
- May 2024, at CoRL 20243840
More
arXiv v1 on 2024-05-09. Published at CoRL 2024 (PMLR volume 270, pages 3705-3728; volume dated 2025-01-12).
The GitHub repository was created on 2024-03-23, before the paper.
- Version
- 0.0.1. There are no tagged releases.411242+1
More
Package simpler_env 0.0.1 with the submodule mani_skill2_real2sim 0.5.3 (pins SAPIEN 2.2.2). No tags or releases. Two branches: main (CPU, ManiSkill2) and maniskill3 (GPU versions of the 4 WidowX tasks).
setup.py files read 2026-10-10. The GitHub tags and releases lists are empty. arXiv has only v1 (2024-05-09); the PMLR camera-ready text differs from it (validity.v4, issues.i9).
main branch · SAPIEN 2.2.2 and ManiSkill2 on CPU. Used for all of the paper's results.3712
maniskill3 branch and ManiSkill3 digital twins · GPU-parallel versions of the 4 WidowX tasks inside ManiSkill3. The ManiSkill docs say up to 60x faster than real-world evaluation and 10x faster than CPU simulation. The branch README says reproducing the original results needs the main branch.332543
Visual Matching and Variant Aggregation · Two official setups. Visual Matching overlays real background images and tunes object and arm textures. Variant Aggregation averages over randomised backgrounds, lighting, distractors and table textures (Google Robot only).137
- Last update
- September 2026. It was only a fix to a software dependency.423312
More
2026-09-30: a dependency pin (transformers<5) merged on main. Last change to tasks or environments: October 2024 (GPU port of the WidowX tasks).
Other commits on main since 2025: 2025-02-25, 2025-03-28 and 2025-12-20 (README CUDA note). The maniskill3 branch was last changed 2024-10-30; ManiSkill2_real2sim was last pushed 2024-10-19.
Setup
- Runs in
- Simulation1
More
Real-to-sim: simulated copies of specific real evaluation setups.
- Simulator
- SAPIEN and ManiSkill211233+1
More
Paper Section V builds on SAPIEN; Section VI-D reproduces the Google Robot Variant Aggregation setup in Isaac Sim (no Isaac Sim code found in the repos). Control runs at 3 Hz for the Google Robot and 5 Hz for WidowX, simulation at about 500 Hz (README). One environment renders 3.5k simulation steps per second on an RTX 4090 at 640 x 512 (paper).
- Robot
- One arm, Arm on wheels1
More
The Google Robot is a mobile manipulator placed at fixed floor positions; its base does not move during a task. The WidowX-250 6-DoF is a fixed single arm.
- Robot model
- Google Robot (RT-1 robot); WidowX-250 6-DoF (BridgeData V2 setup)1
More
The Google Robot uses a head-mounted camera and the WidowX setup a fixed Logitech C920 third-person camera. Neither setup has a wrist camera.
- Setting
- Tabletop Inferred1
More
Tabletop and cabinet-drawer tasks. Real background images are pasted behind the simulated objects ('green screening'). The WidowX eggplant task uses a toy sink.
- Tasks
- 8 task families on 2 robots40137
More
8 task families: 4 for the Google Robot (pick coke can, move near, open/close drawer, open drawer and place apple) and 4 for WidowX (spoon on towel, carrot on plate, stack blocks, eggplant in basket). The code registers 10 environments with sub-variants.
PMLR abstract: eight task families. The README lists 10 prepackaged environments, including a pick-random-object environment that the paper's validation does not use.
- Training data
- No demonstrations ship with SimplerEnv. Policies are trained on real robot data: the RT-1 (Fractal) data for the Google Robot and BridgeData V2 for WidowX.137
More
System identification used 20 trajectories from these datasets (camera-ready Section 4.1). The benchmark does not restrict training data, so users can also train on simulated demonstrations (issues.i3, issues.i8).
- Changes at test
- Start positions change. One setup also changes the visuals. Inferred1
More
Object and robot start positions vary over fixed grids. The Variant Aggregation setup also varies background, lighting, distractors and table texture.
The tasks were chosen to be representative of tasks in the training datasets (paper Section V), so test tasks match training tasks. Visual Matching keeps visuals fixed to the real setup.
Scoring and access
- Scored by
- Success rate1
More
Binary success per episode. WidowX results also report partial 'grasp' success. MMRV and Pearson r are used to validate the simulator, not to score policies.
- Score
- Success rate, reported as three headline averages Inferred11415+2
More
Papers report three headline numbers: the WidowX average over 4 tasks (Visual Matching), the Google Robot Visual Matching average and the Google Robot Variant Aggregation average. Google Robot averages cover 3 task groups in some papers and 4 in others.
The SIMPLER paper itself reports per-task results and correlation metrics, not one average. In the paper, Google Robot Visual Matching results are averaged over four tuned arm colours and Octo results over three seeds (Section VI-A).
- Trials
- 24 to 300 per task, depending on the paper13031+4
More
Real-world grids (paper) · Google Robot: pick coke can 75, move near 60, open/close drawer 54, drawer + apple 27. WidowX: 24 per task. Simulation multiplies these by variants, arm colours and seeds.1
Lead author's advice · Run at least 75 trials per Bridge task (repeat the 24-trial grid 3 times); 25 trials is too noisy.3031
CogACT · Each WidowX task repeated 5 times.13
FASTer · 120 trials per WidowX task instead of the default 24.19
STARE-VLA · 300 evaluation episodes, averaged over 3 seeds; best checkpoint chosen by the same evaluation.24
2026 audit reruns · Each of the 24 grid start states repeated 12 times (288 episodes per task) with the standard step limit (60 steps for stacking).27
- Who runs it
- Each team tests its own model Inferred3937
More
No organiser runs submissions. Each paper runs its own evaluation.
- Error bars
- Not reported Inferred12713+7
More
In the 14 SimplerEnv papers we opened, SimplerEnv success rates are single numbers without intervals (some give intervals for LIBERO only). The SIMPLER paper runs Kruskal-Wallis tests per policy but gives no intervals. The 2026 audit found 62 of 122 claimed WidowX gains cannot be tested from published averages.
- Leaderboard
- None. Scores are only in papers. Inferred393710
More
No leaderboard on the project site or in the README. The closest public compilation is the 2026 audit's tracker CSV of reported numbers.
- Code licence
- MIT (SimplerEnv); Apache-2.0 (ManiSkill2_real2sim submodule with the environments and assets)1112
More
LICENSE files read 2026-10-10.
- Data licence
- MIT for the example evaluation videos on Hugging Face. No training data ships with SimplerEnv; the real-world reference results sit in the MIT-licensed code (metrics.py).4436
More
The RT-1 and BridgeData V2 training datasets carry their own licences (not checked here).
- Asset licence
- Unknown
More
No asset licence list in the paper, README, guide or either repository tree (only the two root LICENSE files). The paper says objects come from Objaverse, from 3D scans of products bought on Amazon, from single-view 3D generation and from manual modelling; Variant Aggregation scenes are modified ReplicaCAD scenes. Objaverse objects each carry their own Creative Commons licence, some non-commercial; ReplicaCAD is CC BY 4.0 (sources s44, s45). Which Objaverse objects were used is not stated.
- Access
- Open. The code is on GitHub.3712
More
Install from GitHub; assets are in the repo; RT-1 checkpoints come from a public Google Cloud bucket and Octo from Hugging Face. Needs an NVIDIA GPU for SAPIEN rendering; no TPU support.
- Commercial use
- Unclear Inferred111245+2
More
Code is MIT and Apache-2.0, which allow commercial use. Some bundled objects come from Objaverse, where individual objects carry their own Creative Commons licences; the Objaverse card lists 25K CC BY-NC and 52K CC BY-NC-SA objects among them. The repo does not say which objects were used or under which terms. Not legal advice.
Sources 46
- 1SIMPLER paper, full text arXiv v1 (Tables I-XIV, Appendix B)Paper · May 2024 · checked 10 Oct 2026
- 2AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World (v2; Section 5.2, Fig. 7, Appendix C-D, Tables 1-3)Paper · Mar 2025 · checked 10 Oct 2026
- 3A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation (Table 2, Appendix A)Paper · Jun 2026 · checked 10 Oct 2026
- 4SIMPLER camera-ready PDF in PMLR (Fig. 4, Tables 1-3, adds OpenVLA-7B)Paper · Jan 2025 · checked 10 Oct 2026
- 5ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (Section III-E, Appendix K, Fig. 25)Paper · Oct 2024 · checked 10 Oct 2026
- 6PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Sections 2 and 5.1)Paper · Dec 2025 · checked 10 Oct 2026
- 7OpenVLA: An Open-Source Vision-Language-Action Model (v1-v3 full text checked for SIMPLER)Paper · Jun 2024 · checked 10 Oct 2026
- 8RobotArena Infinity: Scalable Robot Benchmarking via Real-to-Sim Translation (Sections 5.2-5.3)Paper · Oct 2025 · checked 10 Oct 2026
- 9SimplerEnv issue #78: Poor performance of OpenVLA on BridgeRepository · Mar 2025 · checked 10 Oct 2026
- 10Audit repository: SimplerEnv citation tracker CSV (snapshot used for the paper, 2026-05-21)Repository · Jun 2026 · checked 10 Oct 2026
- 11SimplerEnv LICENSE (MIT)Repository · Mar 2024 · checked 10 Oct 2026
- 12ManiSkill2_real2sim submodule: LICENSE (Apache-2.0), README, setup.py (0.5.3, sapien==2.2.2), asset foldersRepository · Oct 2024 · checked 10 Oct 2026
- 13CogACT (Tables 1-2: SIMPLER Google Robot and WidowX)Paper · Nov 2024 · checked 10 Oct 2026
- 14SpatialVLA (Tables I-II: SimplerEnv)Paper · Jan 2025 · checked 10 Oct 2026
- 15EO-1: An Open Unified Embodied Foundation Model for General Robot Control (Table 4: SimplerEnv)Paper · Aug 2025 · checked 10 Oct 2026
- 16InternVLA-M1 (Tables 1-2: SimplerEnv)Paper · Oct 2025 · checked 10 Oct 2026
- 17X-VLA (Table 2 and Table 12: Simpler)Paper · Oct 2025 · checked 10 Oct 2026
- 18Dexbotic: Open-Source Vision-Language-Action Toolbox (Table 1: SimplerEnv-Bridge)Paper · Oct 2025 · checked 10 Oct 2026
- 19FASTer: Toward Efficient Autoregressive Vision Language Action Modeling (Table 1 and Appendix: Simpler-Bridge)Paper · Dec 2025 · checked 10 Oct 2026
- 20Xiaomi-Robotics-0 (Section 4 and Appendix B: SimplerEnv)Paper · Feb 2026 · checked 10 Oct 2026
- 21SimVLA: A Simple VLA Baseline for Robotic Manipulation (Tables 4-5)Paper · Feb 2026 · checked 10 Oct 2026
- 22CORAL: Scalable Multi-Task Robot Learning via LoRA Experts (Tables II-III)Paper · Mar 2026 · checked 10 Oct 2026
- 23piRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (Table 8, Appendix E.1)Paper · Oct 2025 · checked 10 Oct 2026
- 24STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning VLA Models (Table 1, Appendix A.4-A.5)Paper · Dec 2025 · checked 10 Oct 2026
- 25ManiSkill documentation: Digital Twins (BridgeData v2 evaluation environments)Repository · 2025 · checked 10 Oct 2026
- 26SimplerEnv-OpenVLA (community fork adding OpenVLA support; GitHub API)Repository · Jun 2025 · checked 10 Oct 2026
- 27What Are We Actually Benchmarking in Robot Manipulation? (2026 audit; Sections 4-6, Appendix A.2, A.3, B)Paper · Jun 2026 · checked 10 Oct 2026
- 28Audit repository: statistical-significance counts, paper vs corrected release (commit fe18d73, 2026-10-08)Repository · 8 Oct 2026 · checked 10 Oct 2026
- 29Audit repository: SimplerEnv fixed-grid calibration reruns (per_policy_summary.csv, per_task_summary.csv)Repository · Sep 2026 · checked 10 Oct 2026
- 30SimplerEnv issue #108: Unstable success rates for evaluating OpenVLA (reply by lead author)Repository · Jun 2025 · checked 10 Oct 2026
- 31SimplerEnv issue #125: Different results of configuration evaluation (reply by lead author)Repository · Dec 2025 · checked 10 Oct 2026
- 32SimplerEnv issue #130: Inconsistent Simulation Environment (no reply)Repository · Apr 2026 · checked 10 Oct 2026
- 33SimplerEnv maniskill3 branch README and commitsRepository · Oct 2024 · checked 10 Oct 2026
- 34World-Gymnast: Training Robots with Reinforcement Learning in a World Model (Table 1, Appendix B.2)Paper · Feb 2026 · checked 10 Oct 2026
- 35SimplerEnv issue #77: Source of Real Evaluation Results (reply by lead author)Repository · Mar 2025 · checked 10 Oct 2026
- 36SimplerEnv metrics.py (REAL_PERF and SIMPLER_PERF tables)Repository · May 2024 · checked 10 Oct 2026
- 37SimplerEnv README (main branch)Repository · Dec 2025 · checked 10 Oct 2026
- 38Evaluating Real-World Robot Manipulation Policies in Simulation (arXiv abstract page)Paper · May 2024 · checked 10 Oct 2026
- 39SIMPLER project siteOfficial site · May 2024 · checked 10 Oct 2026
- 40SIMPLER, PMLR proceedings page (CoRL 2024, PMLR 270:3705-3728)Paper · Jan 2025 · checked 10 Oct 2026
- 41SimplerEnv setup.py (simpler_env 0.0.1)Repository · 2024 · checked 10 Oct 2026
- 42SimplerEnv commit history, tags and releases (main branch)Repository · 30 Sep 2026 · checked 10 Oct 2026
- 43SimplerEnv issue #121: Moving ManiSkill2_real2sim tasks to ManiSkill3 (reply by lead author)Repository · Oct 2025 · checked 10 Oct 2026
- 44Hugging Face dataset card: xuanlinli17/simpler-env-eval-example-videos (license: mit)Repository · Apr 2024 · checked 10 Oct 2026
- 45Objaverse dataset card (ODC-By; per-object Creative Commons licences)Repository · 2023 · checked 10 Oct 2026
- 46ReplicaCAD dataset card (CC BY 4.0)Repository · 2023 · checked 10 Oct 2026
Where we searched for missing information
validity (paired sim-vs-real studies of SimplerEnv): SIMPLER arXiv v1 and PMLR camera-ready (all tables); ManiSkill3 paper v1 and v2; AutoEval paper v2 and camera-ready; Wang et al. 2606.10366; PolaRiS v1, v2 and RSS 2026 version (no SIMPLER comparison run); OpenVLA v1-v3 (no mention of SIMPLER, although PolaRiS cites it for SIMPLER's weak correlation); RobotArena Infinity (sim-vs-sim only; its real-world check of one task did not involve SIMPLER); WorldEval 2505.19017 (applied SIMPLER's techniques to RoboTwin, said a direct SIMPLER comparison was not feasible); Scalable Policy Evaluation with Video World Models 2511.11520, WorldGym 2506.00613, Ctrl-World 2510.10125, Veo world simulator 2512.10675, GSWorld 2510.20813, Real-is-Sim 2504.03597, soft-body splat evaluation 2511.04665, RoboArena 2506.18123, SimFoundry 2606.28276, H2RBench 2609.24778 (none pairs SimplerEnv with real results); EchoArena (CVPR 2026 workshop; compares a world model with SimplerEnv, not with real robots); SimplerEnv GitHub issues; web searches (extended) for SimplerEnv sim-real correlation studies.
license_assets: Paper Appendix C, README, ADDING_NEW_ENVS_ROBOTS.md, full git trees of SimplerEnv and ManiSkill2_real2sim (only root LICENSE files), Hugging Face cards for ReplicaCAD and Objaverse.
leaderboard: Project site, README, GitHub repo; the 2026 audit's tracker is a compilation, not an official board.
top_score: 2026 audit tracker and previous-SOTA export (WidowX column) and the 14 papers listed in sources; values read in each paper. Not searched beyond the audit snapshot (2026-05-21) except papers already found.
uncertainty_reported: SimplerEnv tables in CogACT, SpatialVLA, Magma, ThinkAct, EO-1, InternVLA-M1, X-VLA, Dexbotic, FASTer, Xiaomi-Robotics-0, SimVLA, CORAL, piRL and STARE-VLA.
Change history
- Created at full depth from primary sources, starting from the basic entry and the real-eval inventory record. Added the camera-ready validation numbers, the ManiSkill3, AutoEval and Wang et al. comparisons, the audit's corrected significance counts and reruns, verified top scores and adoption counts.
- Published as a full entry.