robomimic

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

How to read this picture

robomimic is a set of five simulated robot-arm tasks with human demonstrations, used to test imitation-learning methods (methods that learn by copying demonstrations). The easier tasks are now near 100% success.123

Sources
Last checked 10 Oct 2026Full entry47 of 60 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation14
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm, Two arms1
Franka Emika Panda
Used by
A common test for imitation-learning methods356+4
1,112 citations
Licence
MIT11
Commercial use: allowed

What a score here does not tell you Inferred

  1. How well a policy will do on a real robot.Only the authors have run the real-robot versions of the tasks. No study compares the same policies in simulation and on real robots.
  2. Which of the top methods is better.Top methods reach 100% on most task variants, so these variants no longer separate them.
  3. How well a policy handles new objects, scenes or instructions.Each task uses the same objects every time, and the tasks have no language instructions.

Comparisons with real robots

Tried on real robots, not compared Inferred1

Details

Real tasks were built to match the simulated ones in look, size and start randomization. BC-RNN on 200 real demos per task, final checkpoint, 30 rollouts: Lift 96.7%, Can 73.3%, Tool Hang 3.3%. Removing image randomization or the wrist camera hurt in both simulation and the real Can task (26.7% and 43.3%), which the authors read as 'study results transfer to real-world settings'. Real policies are different policies trained on real data, so this is not a score comparison.

Our assessment Opinion

The easy tasks no longer separate methods. Look at results on Tool Hang and on the multi-human (MH) data.

Reasoning

On the proficient-human (PH) data, Lift, Can and Square no longer separate strong methods. Differences now show on Tool Hang, on Transport with mixed-quality data, and in low-data settings.

Confidence: high

Compare numbers only when the dataset version and the checkpoint rule match.

Reasoning

Compare robomimic numbers only when the dataset type, observation type, dataset version and checkpoint rule all match. Best-checkpoint numbers are optimistic because the checkpoint is chosen on the test environment.

Confidence: high

robomimic is a small test of learning algorithms. It does not test general robot skill.

Reasoning

A robomimic score says little about language, new objects, new scenes or real robots. It is useful as a small, reproducible test of imitation-learning algorithms, which is how most later papers use it.

Confidence: medium

Known problems 3

  1. Top methods score 100% on most task variants

    Since 2023, top methods have reported a success rate of 1.00 (100%) on most task variants.357

    Details

    Diffusion Policy (2023) reports best-checkpoint success of 1.00 on Lift and Can (PH and MH) and on Square PH and Transport PH with image inputs. AWE (2023) says Diffusion Policy is near-perfect on Lift, Can and Square with 200 proficient demonstrations, so it studies low-data settings instead. DPPO reports RL fine-tuning converging to about 100% on Lift and Can.

  2. Dataset versions and subsets differ between papers

    Papers use different dataset versions, subsets and inputs. Their results are rarely directly comparable.437

    Details

    Three dataset generations exist (mujoco-py offline_study, robosuite 1.4.1, robosuite 1.5.1), and the docs warn results may not match the CoRL 2021 datasets. Image observations are re-rendered locally from raw states. Papers choose PH or MH data, state or image inputs, and four or five tasks. Diffusion Policy's re-run of BC-RNN scored slightly better than the original paper.

  3. Scores depend on which checkpoint is reported

    Scores from the best checkpoint (a saved copy of the model during training) can be far above scores from the last checkpoint. The robomimic study found 10% to 100% lower success when it took the last checkpoint or picked one by validation loss.13

    Details

    The study evaluates every checkpoint in the test environment and reports the best, with no separate validation set. Its own analysis finds that picking by validation loss or taking the last checkpoint gives 10% to 100% lower success. Diffusion Policy prints both numbers: for example 0.68 best vs 0.46 last-10 average on Transport MH (state, CNN) and 0.76 vs 0.47 on Tool Hang PH (image, Transformer).

Details

About

What it is
Fixed tasks and datasets for imitation learning Inferred213
More

The authors call robomimic 'a framework for robot learning from demonstration' that lets researchers 'benchmark tasks and algorithms fairly'. Its five simulated tasks have fixed datasets and a success-rate evaluation; Diffusion Policy calls it 'a large-scale robotic manipulation benchmark'. The library alone would not be a benchmark.

Built by
Stanford University, The University of Texas at Austin112
More

Authors: Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, Roberto Martín-Martín. Part of the ARISE Initiative; development began in the Stanford Vision and Learning Lab in late 2018 (README).

Stanford University · 8 of 10 authors (Stanford Vision and Learning Lab)12

The University of Texas at Austin · Soroush Nasiriany and Yuke Zhu1

Released
August 2021, at CoRL 2021131415+1
More

arXiv v1 2021-08-06; repository created 2021-08-06; CoRL 2021 (PMLR 164).

Version
The library is at 0.5.0 and the datasets are v0.1.1724
More

Library v0.5.0 (released 2025-06-27). Datasets: 'robomimic v0.1' (CoRL 2021), now distributed as robosuite v1.5 versions.

v0.1.0 · Code and paper release dated 2021-08-09 in the README; GitHub release 2021-11-16, 'should be used when trying to reproduce results from the study'217

v0.2.0 (2021-12-17) · Modular observation modalities; MOMART datasets17

v0.3.0 (2023-07-04) · BC-Transformer and IQL; robosuite v1.4 and DeepMind MuJoCo bindings17

v0.4.0 (2025-03-11) · robosuite v1.5 support; datasets moved to Hugging Face17

v0.5.0 (2025-06-27) · Diffusion Policy, multi-dataset training, language-conditioned policies17

Three dataset generations · CoRL 2021 datasets used the mujoco-py 'offline_study' branch of robosuite; v0.3 shipped robosuite 1.4.1 versions; v0.4+ ships robosuite 1.5.1 versions. The docs warn learning results 'may not match exactly'.4

Last update
August 2026. The latest changes are minor fixes.181719
More

Last commit 2026-08-09 (lazy CLIP download). Last release v0.5.0 on 2025-06-27. Datasets moved to the robomimic Hugging Face organisation on 2026-02-05.

Status
Maintained. The last release was in June 2025. Inferred1817
More

No release for over a year; small fixes in 2025-09, 2025-10, 2026-02 and 2026-08.

Setup

Runs in
Simulation Inferred14
More

Scores in later papers come from the five simulated tasks. The study also ran three real-world versions (Lift, Can, Tool Hang) on a Franka arm; their datasets are released, but real evaluation needs the authors' physical setup.

Simulator
robosuite, built on MuJoCo14
More

robosuite on MuJoCo; v1.5.1 recommended since robomimic v0.4. The CoRL 2021 datasets used robosuite's mujoco-py 'offline_study' branch.

Policies act at 20 Hz through an operational space controller.

Robot
One arm, Two arms1
Robot model
Panda arms in simulation and in the real world; Transport uses two.1
Setting
Tabletop1
Tasks
5 simulated and 3 real tasks12
More

8 tasks: 5 simulated (Lift, Can, Square, Transport, Tool Hang) and 3 real (Lift, Can, Tool Hang)

Later papers use the 5 simulated tasks, often only 4 (Tool Hang has no multi-human dataset).

Training data
200 to 300 human demonstrations per task Inferred14
More

3,000 human demonstrations by our arithmetic, plus 5,400 machine-generated trajectories

PH: 200 per task for 5 simulated and 3 real tasks (1,600). MH: 300 per task for 4 tasks (1,200). Paired: 200 (Can). MG: 1,500 (Lift) and 3,900 (Can) from SAC checkpoints.

Proficient-Human (PH) · 200 demos per task from one experienced operator via RoboTurk14

Multi-Human (MH) · 300 demos per task from 6 operators of 3 skill levels (50 each); Lift, Can, Square, Transport14

Machine-Generated (MG) · Rollouts from SAC checkpoints: 1,500 (Lift), 3,900 (Can)4

Paired · 200 Can demos: one success and one failure for each of 100 starts14

Hugging Face size · 26 files, about 6.56 GB (low-dim and raw state files; image versions are generated locally)20

Real datasets · Lift (1.9 GB), Can (5.3 GB), Tool Hang (58 GB) on a Stanford server421

Changes at test
Only the start positions change1
More

Object poses are randomized at the start of each episode, within small regions. Objects and scenes stay the same.

Scoring and access

Scored by
Success rate1
Score
Success rate per task, taken from the best checkpoint saved during training13
More

Success over 50 rollouts, evaluated every few epochs; the study reports the best success over training, averaged over 3 seeds. Results are given per task, per dataset (PH, MH, MG) and per observation type (low-dim or image). There is no standard cross-task average.

Diffusion Policy reports both the best checkpoint and the average of the last 10 checkpoints.

Trials
50 attempts per evaluation, averaged over 3 training runs (seeds)13
More

Study · 50 rollouts every 50 epochs (low-dim) or 20 epochs (image); 3 seeds; best over training1

Study, real robot · Final checkpoint, 30 rollouts1

Diffusion Policy · 3 seeds x 50 initial conditions (150); best checkpoint and average of last 10 checkpoints3

Who runs it
Each team tests its own model Inferred212
Error bars
Sometimes reported Inferred13
More

The study reports mean and standard deviation over 3 seeds. Diffusion Policy gives seed averages without error bars.

Leaderboard
None. Scores are only in papers. Inferred2124
More

No leaderboard on the site, docs or README.

Code licence
MIT11
More

LICENSE: MIT, copyright 2021 Stanford Vision and Learning Lab.

Data licence
Hugging Face dataset card: mit (simulated datasets)20
More

The card says the repository holds 'some of the datasets'. The real-robot datasets on the Stanford server carry no licence statement that we found.

Asset licence
MIT, through robosuite, by our reading Inferred22
More

Task assets ship in robosuite, whose LICENSE is MIT (with an Apache-2.0 notice for partial MuJoCo code). No separate asset licence.

Reading the robosuite root licence as covering its bundled meshes is ours; robosuite has no other licence file.

Access
Open. The data is on Hugging Face.202123+1
More

Simulated datasets on an ungated Hugging Face repository (moved to robomimic/robomimic_datasets on 2026-02-05; the old amandlek/robomimic id redirects). Real datasets on a Stanford server: Lift and Can links answered HTTP 200 on 2026-10-10; the 58 GB Tool Hang link returned HTTP 503 twice.

Commercial use
Allowed Inferred112022
More

Code, simulated datasets and robosuite are MIT. The real-robot datasets have no stated licence, so their use is unclear. Not legal advice.

Sources 23

  1. 1robomimic study paper, full text v2 (Tables 1 to 3, Sections 3 and 4, Appendix E)Paper · Sep 2021 · checked 10 Oct 2026
  2. 2robomimic README (latest updates, description)Repository · Jun 2025 · checked 10 Oct 2026
  3. 3Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (Tables 1 and 2)Paper · Mar 2023 · checked 10 Oct 2026
  4. 4robomimic docs: robomimic v0.1 (CoRL 2021) datasets, versions warning, dataset info, reproduction stepsOfficial site · 2025 · checked 10 Oct 2026
  5. 5Waypoint-Based Imitation Learning for Robotic Manipulation (AWE)Paper · Jul 2023 · checked 10 Oct 2026
  6. 6Consistency Policy: Accelerated Visuomotor Policies via Consistency DistillationPaper · May 2024 · checked 10 Oct 2026
  7. 7Diffusion Policy Policy Optimization (DPPO)Paper · Sep 2024 · checked 10 Oct 2026
  8. 8MimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsPaper · Oct 2023 · checked 10 Oct 2026
  9. 9RoboCasa v0.2 docs: Policy Learning (official code is a robomimic branch)Repository · Apr 2025 · checked 10 Oct 2026
  10. 10LIBERO requirements.txt (robomimic==0.2.0)Repository · 10 Dec 2024 · checked 10 Oct 2026
  11. 11robomimic LICENSE (MIT, 2021 Stanford Vision and Learning Lab)Repository · 2021 · checked 10 Oct 2026
  12. 12robomimic project siteOfficial site · 2023 · checked 10 Oct 2026
  13. 13What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (arXiv abstract page)Paper · Aug 2021 · checked 10 Oct 2026
  14. 14GitHub API: ARISE-Initiative/robomimic (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
  15. 15PMLR volume 164 page (Proceedings of the 5th Conference on Robot Learning, 164:1678-1690)Paper · 2022 · checked 10 Oct 2026
  16. 16robomimic study pageOfficial site · 2021 · checked 10 Oct 2026
  17. 17robomimic GitHub releases (v0.1.0 to v0.5.0) and release notesRepository · 27 Jun 2025 · checked 10 Oct 2026
  18. 18robomimic commit history (master)Repository · 9 Aug 2026 · checked 10 Oct 2026
  19. 19robomimic pull request #294 'update HF_REPO_ID' (datasets moved to a new Hugging Face organisation)Repository · 5 Feb 2026 · checked 10 Oct 2026
  20. 20Hugging Face dataset robomimic/robomimic_datasets: card (license mit) and file listingRepository · 13 Apr 2025 · checked 10 Oct 2026
  21. 21robomimic dataset registry (robomimic/__init__.py: horizons, real-data links on Stanford server)Repository · 5 Feb 2026 · checked 10 Oct 2026
  22. 22robosuite LICENSE (MIT, 2022 Stanford Vision and Learning Lab and UT Robot Perception and Learning Lab)Repository · 2022 · checked 10 Oct 2026
  23. 23robomimic real-robot datasets on the Stanford download server (Lift, Can, Tool Hang)Repository · 2021 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: Study paper full text (Section 4.7, Appendix E.2); study page; docs; Diffusion Policy, DPPO, MimicGen (real experiments use other tasks); evidence papers PolaRiS (2512.16881), SureSim (2510.04354), 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366), Betting for Sim-to-Real (2604.24018), Robot Policy Evaluation for Sim-to-Real Transfer (2508.11117), Active Real-World Factor-Based Evaluation (2607.14439), Beyond Binary Success (2603.13616): no robomimic mention in full text; web searches. No paired study found.

license_data (real datasets): Docs datasets page, README, Hugging Face card, dataset registry: no licence statement for the Stanford-hosted real datasets.

used_by (count): No tracker found; the docs' 'Projects using robomimic' page only links to Google Scholar.

issues (audits): The 2026 audit (2606.04233) does not cover robomimic; web search for robomimic saturation found AWE's statement only.

Change history

  1. Created full entry from primary sources (no basic entry existed). Scope decided: in scope as a benchmark of fixed tasks and datasets.