SimplerEnv

Evaluating Real-World Robot Manipulation Policies in Simulation

How to read this picture

SimplerEnv is a set of simulated copies of two real robot test setups, used to score robot policies (the models that control a robot) in simulation. Its scores agreed with real-robot results for policies from 2023 and 2024, and later checks with newer policies found weaker agreement.123

Sources
Last checked 10 Oct 2026Full entry62 of 77 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Checked
Measured by its authors. Later checks found weaker agreement.145+6
Skill
Handling objects
Robot
One arm, Arm on wheels1
Used by
207 papers10
578 citations
Licence
Unclear1112
Commercial use: unclear

What a score here does not tell you Inferred

  1. How well current vision-language-action (VLA) models will do on real robots.The authors validated it on policies from 2023 and 2024. Later checks with newer policies found weaker agreement.
  2. Whether a small improvement is real.Most claimed improvements cannot be shown to be statistically significant.
  3. How well a policy generalises beyond its training data.Policies trained on data close to the test, or inside the test simulator, can reach top scores.
ChartPublished scores over time
95% AND ABOVE3040506070809010020252026CogACT · 51.3% · 2024-11SpatialVLA · 42.7% · 2025-01EO-1 · 72.7% · 2025-08InternVLA-M1 · 71.7% · 2025-10X-VLA (0.9B) · 95.8% · 2025-10Dexbotic DB-MemVLA · 84.4% · 2025-10FASTer · 87.9% · 2025-12Xiaomi-Robotics-0 · 79.2% · 2026-02SimVLA · 95.8% · 2026-02CORAL (SimVLA) · 97.9% · 2026-03piRL on pi0 · 86.7% · 2025-10STARE-VLA (IPI) · 98% · 2025-12CogACT 51.3%CORAL 97.9%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.131415+9

Comparisons with real robots

StudyResultWhat was comparedDone by
SIMPLER paper: Google Robot, Visual Matching
May 2024
Pearson r = 0.924 and MMRV = 0.056 (mean of 3 task groups)1The same 6 checkpoints (3 RT-1 checkpoints, RT-1-X, RT-2-X and Octo-Base, so 4 distinct models) were scored in simulation and on the real Google Robot, on the same tasks and trial grids. Agreement was measured with Pearson r and MMRV (a measure of how often two rankings disagree). The study’s authors described the result as “strong”.The benchmark’s authors
SIMPLER paper: Google Robot, Variant Aggregation
May 2024
Pearson r = 0.778 and MMRV = 0.143 (mean of 3 task groups)1The same 6 checkpoints were scored in simulation with randomised backgrounds, lighting, distractors and table textures, and compared with the same real results.The benchmark’s authors
SIMPLER paper: WidowX and BridgeData V2
May 2024
Pearson r = 0.890 and MMRV = 0.014 (mean over 4 tasks, for success and grasp)13 policies (RT-1-X, Octo-Base and Octo-Small, so 2 model families) were scored in simulation and on a real WidowX on the same 4 tasks, with 24 real trials per task. The study’s authors described the result as “strong”.The benchmark’s authors
SIMPLER camera-ready: Google Robot with OpenVLA added
Jan 2025
Pearson r = 0.929 and MMRV = 0.049 (Fig. 4)47 checkpoints (the 6 above plus OpenVLA-7B, so 5 distinct models) were scored in simulation (Visual Matching) and on the real Google Robot, on the same tasks. The study’s authors described the result as “strong”.The benchmark’s authors
SIMPLER paper: Isaac Sim re-implementation
May 2024
Pearson r = 0.919 and MMRV = 0.058. In SAPIEN, r = 0.923 and MMRV = 0.082.15 checkpoints (3 RT-1 checkpoints, RT-1-X and Octo-Base) were scored on the pick coke can and move near tasks in Variant Aggregation scenes rebuilt in Isaac Sim, and compared with the same real results.The benchmark’s authors
SIMPLER paper: sensitivity to visual shifts
May 2024
Pearson r = 0.831 and MMRV = 0.000 without augmentation. Pearson r = 0.970 and MMRV = 0.016 with augmentation.12 RT-1 policies, with and without image augmentation, were tested under 5 visual shifts. Their drop in success in simulation was compared with earlier real-world tests by Xie et al. The study’s authors described the result as “accurately reflect”.The benchmark’s authors
ManiSkill3 GPU port
Oct 2024
Pearson r = 0.9284 and MMRV = 0.014753 policies (Octo-Base, Octo-Small and RT-1-X) were run on the GPU version of the 4 WidowX tasks and compared with the SIMPLER paper's real results. Success and grasp rates were pooled into one plot. The study’s authors described the result as “close to that of the original paper”.The benchmark’s authors
AutoEval paper (UC Berkeley)
Mar 2025
Mean Pearson r = 0.548 and MMRV = 0.207 over 4 tasks (our computation from the paper's tables)26 policies (OpenVLA, Octo, Open-pi0, MiniVLA, SuSIE and SuSIE's low-level policy, so 5 distinct models) were run 50 times per task in SIMPLER and in human-run real evaluations on matching WidowX cells. Of the 4 tasks, only eggplant-in-basket is an official SimplerEnv task. The study’s authors described the result as “policy dependent”.The benchmark’s authors
Wang et al. (Tsinghua, Shanghai Qi Zhi)
Jun 2026
Mean Spearman correlation 0.400, Pearson r = 0.402 and MMRV = 0.12835 VLA policies (pi0, pi0-FAST, pi0.5, GR00T N1.6 and GR00T N1.7) were run on 7 tabletop tasks rebuilt in SIMPLER's SAPIEN setting and matched to real tasks on DROID hardware, under changes to vision, layout and language. Each policy had 20 simulated and 5 real test runs per task and change.An independent group

Benchmarks built on SimplerEnv: ManiSkill3 BridgeData v2 digital twins, SimplerEnv-OpenVLA, AutoEval SIMPLER scenes, Audit stack-task variants.25262+1

Our assessment Opinion

Scores agreed with real-robot results for older policies. The evidence for current VLA models is weak.

Reasoning

A SimplerEnv score is good evidence of real-robot ranking only for policies like those its authors tested (RT-1, RT-2-X, Octo, OpenVLA) on these two setups. For current VLA models the evidence is weak. The two later studies that included newer policies found r = 0.548 and r = 0.402, both mostly on tasks outside the official set.

Confidence: medium

Check the protocol before comparing numbers across papers.

Reasoning

Do not compare WidowX numbers across papers without checking trials, step limits, training data and which baseline values were copied. Reported numbers for the same model differ by up to 42 points, and independent reruns came in lower than published.

Confidence: high

Top WidowX scores say little about real-world generalisation.

Reasoning

Above about 95% on WidowX, scores no longer separate methods, and such scores can be reached by training close to the test or inside the simulator. Scores on the Google Robot settings are further from 100%.

Confidence: medium

Its method for checking a simulator against real robots has become the common standard.

Reasoning

SimplerEnv's method for checking a simulator against real robots has become the common standard. It scores the same policies in simulation and on real robots and compares the results with Pearson r (a correlation) and MMRV (a measure of how often two rankings disagree). AutoEval, PolaRiS and later studies report the same two numbers.

Confidence: medium

Known problems 10

  1. Most claimed improvements are not shown to be statistically significant

    Only about 1 in 5 claimed improvements on WidowX can be shown to be statistically significant from the published scores.2728

    Details

    The 2026 audit checked 122 previous-best-to-new comparisons on the WidowX protocol, using only published scores. Its paper classes 19.7% (24) as provably significant at the 5% level, 50.8% (62) as inconclusive, and the rest as no improvement or provably not significant. On 2026-10-08 the audit authors released corrected classifications: 27 of 122 provably significant (22.1% by our arithmetic), 62 inconclusive, 16 provably not significant and 17 no improvement.

  2. Independent reruns score lower than the published numbers

    When a 2026 audit reran five published policies, they scored 2.6 to 23.4 points below the published numbers.292713+4

    Details

    The 2026 audit reran five published WidowX policies on the official grid with the standard step limits (each of the 24 start states repeated 12 times; 1,152 episodes per policy). Results: CogACT-Base 48.7%, SpatialVLA 37.5%, InternVLA-M1 61.2%, X-VLA 72.4% and Dexbotic DB-MemVLA 64.7%. The papers report 51.3%, 42.7% (34.4% without fine-tuning), 71.7%, 95.8% and 84.4%. On the stack task alone, X-VLA reports 95.8% and the audit measured 59.7% (172/288). The audit says several of these policies report under longer, easier episode step limits.

  3. Training data placed next to the test can match top scores

    A 2026 audit trained small models (22 million parameters) on 120 scripted demonstrations per task, recorded in simulation next to the test positions. They scored 91 of 96 trials, against 92 of 96 reported by X-VLA.27

    Details

    SimplerEnv tests in simulation, while its standard training data are real BridgeData V2 demonstrations, so a high score reads as transfer across that gap. The benchmark does not restrict training data. The 2026 audit trained a separate 22M-parameter policy per task, with no robotics pretraining, on 120 scripted demonstrations recorded in simulation next to the official test grid. Together the four policies scored 91/96 (94.8%), against 92/96 (95.8%) reported by X-VLA. The audit calls this an existence result: the score alone cannot tell crossing the gap from removing it. The same audit found no shortcut of the LIBERO kind: a small probe trained on the standard Bridge data scored 0.0%.

  4. Results vary between runs and machines

    One user saw the same model score between 16% and 40% on one task across runs.303127+2

    Details

    The lead author advises at least 75 trials per Bridge task because the 24-trial grid is too noisy. A user saw 16% to 40% success for the same OpenVLA model on one task across runs. Another user's repeated RT-1-X runs differed from the paper and from each other, although RT-1-X outputs are deterministic; the lead author attributes this to simulator nondeterminism. The 2026 audit found SimplerEnv rollouts with CogACT diverged from step 0 when only the CPU or only the GPU was changed. A 2026 issue reports the task and object positions changing between runs with a fixed seed (no reply by 2026-10-10). The GPU (ManiSkill3) version is a separate implementation; its README says reproducing the paper needs the main branch.

  5. Small changes seen in the training data lower scores on the stack task

    When the cubes started on support blocks, X-VLA's success on the stack task fell from 59.7% to 31.3%.27

    Details

    The 2026 audit changed the WidowX stack task in three ways that also occur in the BridgeData V2 training data: reversed colour order in the instruction, cubes starting on support blocks, and random cube and arm start poses (288 episodes per policy each). Stacked supports cut X-VLA from 172/288 to 90/288 (28.47 points, 95% CI 20.46 to 35.96) and Dexbotic from 129/288 to 79/288 (17.36 points). The reversed instruction cut CogACT by 11.11 points and InternVLA-M1 by 9.38. The audit reads these drops as overfitting to the fixed test grid.

  6. Top WidowX scores are close to 100%

    By early 2026, top WidowX averages reached 95.8% to 97.9%.172122+2

    Details

    Since October 2025, X-VLA (95.8%), SimVLA (95.8%) and CORAL (97.9%) report WidowX averages above 95% without RL in the simulator, and STARE-VLA reports 98.0% after RL inside SimplerEnv. Several tasks have reported scores of 100%. Google Robot results are lower: the highest we found are 85.5% (Visual Matching) and 84.9% (Variant Aggregation).

  7. The same model gets very different scores in different papers

    Different papers report pi0's WidowX average as anywhere from 27.1% to 69.2%.152019+5

    Details

    WidowX averages reported for pi0 by other papers: 69.2% (EO-1, Xiaomi-Robotics-0), 66.7% (FASTer), 40.1% (SimVLA), 27.8% (X-VLA) and 27.1% (InternVLA-M1). For OpenVLA: 1.0% (SpatialVLA, EO-1), 4.2% (CogACT, InternVLA-M1), 7.8% (SimVLA), 8.3% (X-VLA) and 29.5% (FASTer). Google Robot averages cover 3 task groups in some papers and 4 in others: SpatialVLA reports its own Visual Matching score as 75.1% over 3 tasks, while EO-1 lists SpatialVLA at 55.3% over 4. Some tables do not add up: InternVLA-M1 lists pi0-FAST per-task scores that average 32.1% next to an average of 48.3%; SimVLA's WidowX table gives Octo-Base per-task scores of 12.5 / 8.3 / 0.0 / 43.1 with an average of 31.3. Xiaomi-Robotics-0 (2026-02) calls its 79.2% WidowX average the best result, while its comparison table leaves out X-VLA's 95.8% from 2025-10.

  8. Some top scores come from training inside the test simulator

    Reinforcement learning (RL) inside the simulator raised pi0 from 67.2% to 86.7%.232434

    Details

    piRL raised pi0 from 67.2% to 86.7% on the WidowX tasks with RL inside the simulator. STARE-VLA reports 98.0% after RL fine-tuning in SimplerEnv and reports the best checkpoint under the same evaluation it publishes. World-Gymnast trained RL baselines in SIMPLER copies of the AutoEval scenes and found they transferred poorly to the real cells: for example 34% versus 58% for its world-model method on opening the drawer.

  9. The real-robot check covered older policies on two setups

    The authors' real-robot check covered policies from the RT-1 era on two setups. The policies most reported today were not part of it. Inferred1435+1

    Details

    The authors' validation covers RT-1 checkpoints, RT-1-X, RT-2-X, Octo and, in the camera-ready, OpenVLA on the Google Robot, and RT-1-X and Octo on WidowX. Each per-task correlation rests on 3 to 7 points. The policies most reported today (pi0, pi0.5, GR00T, X-VLA) were not part of it. The real Google Robot evaluations were run by Google staff, so outside groups cannot repeat that comparison (our inference); the WidowX setup can be rebuilt. Visual Matching pastes real images behind a fixed camera view, so it cannot serve policies that use wrist cameras, which most current generalist policies do (PolaRiS paper).

  10. The paper's own tables and versions disagree

    Different tables and versions of the SimplerEnv paper give different correlation numbers.1436

    Details

    In arXiv v1, Table I gives the Google Robot drawer task r 0.942 and MMRV 0.027 under Visual Matching, while Table IV and Fig. 6 give 0.915 and 0.055; the headline mean (r 0.924, MMRV 0.056) uses Table I. Fig. 8 gives MMRV 0.016 for the augmented RT-1 policy where Table VI gives 0.041. The camera-ready adds OpenVLA-7B to the Google Robot study and changes the per-task numbers (drawer r 0.823), but keeps the 6-checkpoint means in its Table 1. For put-eggplant-in-basket, Appendix B states 24 real trials, the real values imply 30, and the repo's metrics.py lists 0.250 and 0.400 for Octo-Base and Octo-Small where the paper has 0.233 and 0.433.

Details

About

What it is
Benchmark Inferred137
More

The paper presents a suite of simulated evaluation environments with fixed tasks, trial grids and success checks; papers report success rates on it as a benchmark. Classification by the Atlas.

Built by
UC San Diego, Stanford University, UC Berkeley, Google DeepMind38139
More

16 authors. Equal contribution: Xuanlin Li, Kyle Hsu, Jiayuan Gu. Co-advising: Hao Su, Quan Vuong, Ted Xiao.

UC San Diego · Hao Su's lab; lead authors and the SAPIEN/ManiSkill code base.1

Stanford University · Kyle Hsu, Chelsea Finn, Jiajun Wu and others.1

UC Berkeley · Sergey Levine's lab; WidowX real-world evaluations.1

Google DeepMind · RT-series policies; the real Google Robot evaluations were run by Google staff, per the lead author in issue #77.135

Released
May 2024, at CoRL 20243840
More

arXiv v1 on 2024-05-09. Published at CoRL 2024 (PMLR volume 270, pages 3705-3728; volume dated 2025-01-12).

The GitHub repository was created on 2024-03-23, before the paper.

Version
0.0.1. There are no tagged releases.411242+1
More

Package simpler_env 0.0.1 with the submodule mani_skill2_real2sim 0.5.3 (pins SAPIEN 2.2.2). No tags or releases. Two branches: main (CPU, ManiSkill2) and maniskill3 (GPU versions of the 4 WidowX tasks).

setup.py files read 2026-10-10. The GitHub tags and releases lists are empty. arXiv has only v1 (2024-05-09); the PMLR camera-ready text differs from it (validity.v4, issues.i9).

main branch · SAPIEN 2.2.2 and ManiSkill2 on CPU. Used for all of the paper's results.3712

maniskill3 branch and ManiSkill3 digital twins · GPU-parallel versions of the 4 WidowX tasks inside ManiSkill3. The ManiSkill docs say up to 60x faster than real-world evaluation and 10x faster than CPU simulation. The branch README says reproducing the original results needs the main branch.332543

Visual Matching and Variant Aggregation · Two official setups. Visual Matching overlays real background images and tunes object and arm textures. Variant Aggregation averages over randomised backgrounds, lighting, distractors and table textures (Google Robot only).137

Last update
September 2026. It was only a fix to a software dependency.423312
More

2026-09-30: a dependency pin (transformers<5) merged on main. Last change to tasks or environments: October 2024 (GPU port of the WidowX tasks).

Other commits on main since 2025: 2025-02-25, 2025-03-28 and 2025-12-20 (README CUDA note). The maniskill3 branch was last changed 2024-10-30; ManiSkill2_real2sim was last pushed 2024-10-19.

Status
Only fixes since 2024. Use is very active. Inferred4210
More

Occasional dependency fixes; no new tasks since 2024. Use is very active (facts.used_by).

33 open issues on 2026-10-10. The lead author still answers issues (latest reply 2026-09-25).

Setup

Runs in
Simulation1
More

Real-to-sim: simulated copies of specific real evaluation setups.

Simulator
SAPIEN and ManiSkill211233+1
More

Paper Section V builds on SAPIEN; Section VI-D reproduces the Google Robot Variant Aggregation setup in Isaac Sim (no Isaac Sim code found in the repos). Control runs at 3 Hz for the Google Robot and 5 Hz for WidowX, simulation at about 500 Hz (README). One environment renders 3.5k simulation steps per second on an RTX 4090 at 640 x 512 (paper).

Robot
One arm, Arm on wheels1
More

The Google Robot is a mobile manipulator placed at fixed floor positions; its base does not move during a task. The WidowX-250 6-DoF is a fixed single arm.

Robot model
Google Robot (RT-1 robot); WidowX-250 6-DoF (BridgeData V2 setup)1
More

The Google Robot uses a head-mounted camera and the WidowX setup a fixed Logitech C920 third-person camera. Neither setup has a wrist camera.

Setting
Tabletop Inferred1
More

Tabletop and cabinet-drawer tasks. Real background images are pasted behind the simulated objects ('green screening'). The WidowX eggplant task uses a toy sink.

Tasks
8 task families on 2 robots40137
More

8 task families: 4 for the Google Robot (pick coke can, move near, open/close drawer, open drawer and place apple) and 4 for WidowX (spoon on towel, carrot on plate, stack blocks, eggplant in basket). The code registers 10 environments with sub-variants.

PMLR abstract: eight task families. The README lists 10 prepackaged environments, including a pick-random-object environment that the paper's validation does not use.

Training data
No demonstrations ship with SimplerEnv. Policies are trained on real robot data: the RT-1 (Fractal) data for the Google Robot and BridgeData V2 for WidowX.137
More

System identification used 20 trajectories from these datasets (camera-ready Section 4.1). The benchmark does not restrict training data, so users can also train on simulated demonstrations (issues.i3, issues.i8).

Changes at test
Start positions change. One setup also changes the visuals. Inferred1
More

Object and robot start positions vary over fixed grids. The Variant Aggregation setup also varies background, lighting, distractors and table texture.

The tasks were chosen to be representative of tasks in the training datasets (paper Section V), so test tasks match training tasks. Visual Matching keeps visuals fixed to the real setup.

Scoring and access

Scored by
Success rate1
More

Binary success per episode. WidowX results also report partial 'grasp' success. MMRV and Pearson r are used to validate the simulator, not to score policies.

Score
Success rate, reported as three headline averages Inferred11415+2
More

Papers report three headline numbers: the WidowX average over 4 tasks (Visual Matching), the Google Robot Visual Matching average and the Google Robot Variant Aggregation average. Google Robot averages cover 3 task groups in some papers and 4 in others.

The SIMPLER paper itself reports per-task results and correlation metrics, not one average. In the paper, Google Robot Visual Matching results are averaged over four tuned arm colours and Octo results over three seeds (Section VI-A).

Trials
24 to 300 per task, depending on the paper13031+4
More

Real-world grids (paper) · Google Robot: pick coke can 75, move near 60, open/close drawer 54, drawer + apple 27. WidowX: 24 per task. Simulation multiplies these by variants, arm colours and seeds.1

Lead author's advice · Run at least 75 trials per Bridge task (repeat the 24-trial grid 3 times); 25 trials is too noisy.3031

CogACT · Each WidowX task repeated 5 times.13

FASTer · 120 trials per WidowX task instead of the default 24.19

STARE-VLA · 300 evaluation episodes, averaged over 3 seeds; best checkpoint chosen by the same evaluation.24

2026 audit reruns · Each of the 24 grid start states repeated 12 times (288 episodes per task) with the standard step limit (60 steps for stacking).27

Who runs it
Each team tests its own model Inferred3937
More

No organiser runs submissions. Each paper runs its own evaluation.

Error bars
Not reported Inferred12713+7
More

In the 14 SimplerEnv papers we opened, SimplerEnv success rates are single numbers without intervals (some give intervals for LIBERO only). The SIMPLER paper runs Kruskal-Wallis tests per policy but gives no intervals. The 2026 audit found 62 of 122 claimed WidowX gains cannot be tested from published averages.

Leaderboard
None. Scores are only in papers. Inferred393710
More

No leaderboard on the project site or in the README. The closest public compilation is the 2026 audit's tracker CSV of reported numbers.

Code licence
MIT (SimplerEnv); Apache-2.0 (ManiSkill2_real2sim submodule with the environments and assets)1112
More

LICENSE files read 2026-10-10.

Data licence
MIT for the example evaluation videos on Hugging Face. No training data ships with SimplerEnv; the real-world reference results sit in the MIT-licensed code (metrics.py).4436
More

The RT-1 and BridgeData V2 training datasets carry their own licences (not checked here).

Asset licence
Unknown
More

No asset licence list in the paper, README, guide or either repository tree (only the two root LICENSE files). The paper says objects come from Objaverse, from 3D scans of products bought on Amazon, from single-view 3D generation and from manual modelling; Variant Aggregation scenes are modified ReplicaCAD scenes. Objaverse objects each carry their own Creative Commons licence, some non-commercial; ReplicaCAD is CC BY 4.0 (sources s44, s45). Which Objaverse objects were used is not stated.

Access
Open. The code is on GitHub.3712
More

Install from GitHub; assets are in the repo; RT-1 checkpoints come from a public Google Cloud bucket and Octo from Hugging Face. Needs an NVIDIA GPU for SAPIEN rendering; no TPU support.

Commercial use
Unclear Inferred111245+2
More

Code is MIT and Apache-2.0, which allow commercial use. Some bundled objects come from Objaverse, where individual objects carry their own Creative Commons licences; the Objaverse card lists 25K CC BY-NC and 52K CC BY-NC-SA objects among them. The repo does not say which objects were used or under which terms. Not legal advice.

Sources 46

  1. 1SIMPLER paper, full text arXiv v1 (Tables I-XIV, Appendix B)Paper · May 2024 · checked 10 Oct 2026
  2. 2AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World (v2; Section 5.2, Fig. 7, Appendix C-D, Tables 1-3)Paper · Mar 2025 · checked 10 Oct 2026
  3. 3A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation (Table 2, Appendix A)Paper · Jun 2026 · checked 10 Oct 2026
  4. 4SIMPLER camera-ready PDF in PMLR (Fig. 4, Tables 1-3, adds OpenVLA-7B)Paper · Jan 2025 · checked 10 Oct 2026
  5. 5ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (Section III-E, Appendix K, Fig. 25)Paper · Oct 2024 · checked 10 Oct 2026
  6. 6PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Sections 2 and 5.1)Paper · Dec 2025 · checked 10 Oct 2026
  7. 7OpenVLA: An Open-Source Vision-Language-Action Model (v1-v3 full text checked for SIMPLER)Paper · Jun 2024 · checked 10 Oct 2026
  8. 8RobotArena Infinity: Scalable Robot Benchmarking via Real-to-Sim Translation (Sections 5.2-5.3)Paper · Oct 2025 · checked 10 Oct 2026
  9. 9SimplerEnv issue #78: Poor performance of OpenVLA on BridgeRepository · Mar 2025 · checked 10 Oct 2026
  10. 10Audit repository: SimplerEnv citation tracker CSV (snapshot used for the paper, 2026-05-21)Repository · Jun 2026 · checked 10 Oct 2026
  11. 11SimplerEnv LICENSE (MIT)Repository · Mar 2024 · checked 10 Oct 2026
  12. 12ManiSkill2_real2sim submodule: LICENSE (Apache-2.0), README, setup.py (0.5.3, sapien==2.2.2), asset foldersRepository · Oct 2024 · checked 10 Oct 2026
  13. 13CogACT (Tables 1-2: SIMPLER Google Robot and WidowX)Paper · Nov 2024 · checked 10 Oct 2026
  14. 14SpatialVLA (Tables I-II: SimplerEnv)Paper · Jan 2025 · checked 10 Oct 2026
  15. 15EO-1: An Open Unified Embodied Foundation Model for General Robot Control (Table 4: SimplerEnv)Paper · Aug 2025 · checked 10 Oct 2026
  16. 16InternVLA-M1 (Tables 1-2: SimplerEnv)Paper · Oct 2025 · checked 10 Oct 2026
  17. 17X-VLA (Table 2 and Table 12: Simpler)Paper · Oct 2025 · checked 10 Oct 2026
  18. 18Dexbotic: Open-Source Vision-Language-Action Toolbox (Table 1: SimplerEnv-Bridge)Paper · Oct 2025 · checked 10 Oct 2026
  19. 19FASTer: Toward Efficient Autoregressive Vision Language Action Modeling (Table 1 and Appendix: Simpler-Bridge)Paper · Dec 2025 · checked 10 Oct 2026
  20. 20Xiaomi-Robotics-0 (Section 4 and Appendix B: SimplerEnv)Paper · Feb 2026 · checked 10 Oct 2026
  21. 21SimVLA: A Simple VLA Baseline for Robotic Manipulation (Tables 4-5)Paper · Feb 2026 · checked 10 Oct 2026
  22. 22CORAL: Scalable Multi-Task Robot Learning via LoRA Experts (Tables II-III)Paper · Mar 2026 · checked 10 Oct 2026
  23. 23piRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (Table 8, Appendix E.1)Paper · Oct 2025 · checked 10 Oct 2026
  24. 24STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning VLA Models (Table 1, Appendix A.4-A.5)Paper · Dec 2025 · checked 10 Oct 2026
  25. 25ManiSkill documentation: Digital Twins (BridgeData v2 evaluation environments)Repository · 2025 · checked 10 Oct 2026
  26. 26SimplerEnv-OpenVLA (community fork adding OpenVLA support; GitHub API)Repository · Jun 2025 · checked 10 Oct 2026
  27. 27What Are We Actually Benchmarking in Robot Manipulation? (2026 audit; Sections 4-6, Appendix A.2, A.3, B)Paper · Jun 2026 · checked 10 Oct 2026
  28. 28Audit repository: statistical-significance counts, paper vs corrected release (commit fe18d73, 2026-10-08)Repository · 8 Oct 2026 · checked 10 Oct 2026
  29. 29Audit repository: SimplerEnv fixed-grid calibration reruns (per_policy_summary.csv, per_task_summary.csv)Repository · Sep 2026 · checked 10 Oct 2026
  30. 30SimplerEnv issue #108: Unstable success rates for evaluating OpenVLA (reply by lead author)Repository · Jun 2025 · checked 10 Oct 2026
  31. 31SimplerEnv issue #125: Different results of configuration evaluation (reply by lead author)Repository · Dec 2025 · checked 10 Oct 2026
  32. 32SimplerEnv issue #130: Inconsistent Simulation Environment (no reply)Repository · Apr 2026 · checked 10 Oct 2026
  33. 33SimplerEnv maniskill3 branch README and commitsRepository · Oct 2024 · checked 10 Oct 2026
  34. 34World-Gymnast: Training Robots with Reinforcement Learning in a World Model (Table 1, Appendix B.2)Paper · Feb 2026 · checked 10 Oct 2026
  35. 35SimplerEnv issue #77: Source of Real Evaluation Results (reply by lead author)Repository · Mar 2025 · checked 10 Oct 2026
  36. 36SimplerEnv metrics.py (REAL_PERF and SIMPLER_PERF tables)Repository · May 2024 · checked 10 Oct 2026
  37. 37SimplerEnv README (main branch)Repository · Dec 2025 · checked 10 Oct 2026
  38. 38Evaluating Real-World Robot Manipulation Policies in Simulation (arXiv abstract page)Paper · May 2024 · checked 10 Oct 2026
  39. 39SIMPLER project siteOfficial site · May 2024 · checked 10 Oct 2026
  40. 40SIMPLER, PMLR proceedings page (CoRL 2024, PMLR 270:3705-3728)Paper · Jan 2025 · checked 10 Oct 2026
  41. 41SimplerEnv setup.py (simpler_env 0.0.1)Repository · 2024 · checked 10 Oct 2026
  42. 42SimplerEnv commit history, tags and releases (main branch)Repository · 30 Sep 2026 · checked 10 Oct 2026
  43. 43SimplerEnv issue #121: Moving ManiSkill2_real2sim tasks to ManiSkill3 (reply by lead author)Repository · Oct 2025 · checked 10 Oct 2026
  44. 44Hugging Face dataset card: xuanlinli17/simpler-env-eval-example-videos (license: mit)Repository · Apr 2024 · checked 10 Oct 2026
  45. 45Objaverse dataset card (ODC-By; per-object Creative Commons licences)Repository · 2023 · checked 10 Oct 2026
  46. 46ReplicaCAD dataset card (CC BY 4.0)Repository · 2023 · checked 10 Oct 2026
Where we searched for missing information

validity (paired sim-vs-real studies of SimplerEnv): SIMPLER arXiv v1 and PMLR camera-ready (all tables); ManiSkill3 paper v1 and v2; AutoEval paper v2 and camera-ready; Wang et al. 2606.10366; PolaRiS v1, v2 and RSS 2026 version (no SIMPLER comparison run); OpenVLA v1-v3 (no mention of SIMPLER, although PolaRiS cites it for SIMPLER's weak correlation); RobotArena Infinity (sim-vs-sim only; its real-world check of one task did not involve SIMPLER); WorldEval 2505.19017 (applied SIMPLER's techniques to RoboTwin, said a direct SIMPLER comparison was not feasible); Scalable Policy Evaluation with Video World Models 2511.11520, WorldGym 2506.00613, Ctrl-World 2510.10125, Veo world simulator 2512.10675, GSWorld 2510.20813, Real-is-Sim 2504.03597, soft-body splat evaluation 2511.04665, RoboArena 2506.18123, SimFoundry 2606.28276, H2RBench 2609.24778 (none pairs SimplerEnv with real results); EchoArena (CVPR 2026 workshop; compares a world model with SimplerEnv, not with real robots); SimplerEnv GitHub issues; web searches (extended) for SimplerEnv sim-real correlation studies.

license_assets: Paper Appendix C, README, ADDING_NEW_ENVS_ROBOTS.md, full git trees of SimplerEnv and ManiSkill2_real2sim (only root LICENSE files), Hugging Face cards for ReplicaCAD and Objaverse.

leaderboard: Project site, README, GitHub repo; the 2026 audit's tracker is a compilation, not an official board.

top_score: 2026 audit tracker and previous-SOTA export (WidowX column) and the 14 papers listed in sources; values read in each paper. Not searched beyond the audit snapshot (2026-05-21) except papers already found.

uncertainty_reported: SimplerEnv tables in CogACT, SpatialVLA, Magma, ThinkAct, EO-1, InternVLA-M1, X-VLA, Dexbotic, FASTer, Xiaomi-Robotics-0, SimVLA, CORAL, piRL and STARE-VLA.

Change history

  1. Created at full depth from primary sources, starting from the basic entry and the real-eval inventory record. Added the camera-ready validation numbers, the ManiSkill3, AutoEval and Wang et al. comparisons, the audit's corrected significance counts and reruns, verified top scores and adoption counts.
  2. Published as a full entry.