robomimic
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
robomimic is a set of five simulated robot-arm tasks with human demonstrations, used to test imitation-learning methods (methods that learn by copying demonstrations). The easier tasks are now near 100% success.123
What a score here does not tell you Inferred
- How well a policy will do on a real robot.Only the authors have run the real-robot versions of the tasks. No study compares the same policies in simulation and on real robots.
- Which of the top methods is better.Top methods reach 100% on most task variants, so these variants no longer separate them.
- How well a policy handles new objects, scenes or instructions.Each task uses the same objects every time, and the tasks have no language instructions.
Comparisons with real robots
Tried on real robots, not compared Inferred1
Details
Real tasks were built to match the simulated ones in look, size and start randomization. BC-RNN on 200 real demos per task, final checkpoint, 30 rollouts: Lift 96.7%, Can 73.3%, Tool Hang 3.3%. Removing image randomization or the wrist camera hurt in both simulation and the real Can task (26.7% and 43.3%), which the authors read as 'study results transfer to real-world settings'. Real policies are different policies trained on real data, so this is not a score comparison.
Our assessment Opinion
The easy tasks no longer separate methods. Look at results on Tool Hang and on the multi-human (MH) data.
Reasoning
On the proficient-human (PH) data, Lift, Can and Square no longer separate strong methods. Differences now show on Tool Hang, on Transport with mixed-quality data, and in low-data settings.
Confidence: high
Compare numbers only when the dataset version and the checkpoint rule match.
Reasoning
Compare robomimic numbers only when the dataset type, observation type, dataset version and checkpoint rule all match. Best-checkpoint numbers are optimistic because the checkpoint is chosen on the test environment.
Confidence: high
robomimic is a small test of learning algorithms. It does not test general robot skill.
Reasoning
A robomimic score says little about language, new objects, new scenes or real robots. It is useful as a small, reproducible test of imitation-learning algorithms, which is how most later papers use it.
Confidence: medium
Known problems 3
Top methods score 100% on most task variants
Since 2023, top methods have reported a success rate of 1.00 (100%) on most task variants.357
Details
Diffusion Policy (2023) reports best-checkpoint success of 1.00 on Lift and Can (PH and MH) and on Square PH and Transport PH with image inputs. AWE (2023) says Diffusion Policy is near-perfect on Lift, Can and Square with 200 proficient demonstrations, so it studies low-data settings instead. DPPO reports RL fine-tuning converging to about 100% on Lift and Can.
Dataset versions and subsets differ between papers
Papers use different dataset versions, subsets and inputs. Their results are rarely directly comparable.437
Details
Three dataset generations exist (mujoco-py offline_study, robosuite 1.4.1, robosuite 1.5.1), and the docs warn results may not match the CoRL 2021 datasets. Image observations are re-rendered locally from raw states. Papers choose PH or MH data, state or image inputs, and four or five tasks. Diffusion Policy's re-run of BC-RNN scored slightly better than the original paper.
Scores depend on which checkpoint is reported
Scores from the best checkpoint (a saved copy of the model during training) can be far above scores from the last checkpoint. The robomimic study found 10% to 100% lower success when it took the last checkpoint or picked one by validation loss.13
Details
The study evaluates every checkpoint in the test environment and reports the best, with no separate validation set. Its own analysis finds that picking by validation loss or taking the last checkpoint gives 10% to 100% lower success. Diffusion Policy prints both numbers: for example 0.68 best vs 0.46 last-10 average on Transport MH (state, CNN) and 0.76 vs 0.47 on Tool Hang PH (image, Transformer).
Details
About
- What it is
- Fixed tasks and datasets for imitation learning Inferred213
More
The authors call robomimic 'a framework for robot learning from demonstration' that lets researchers 'benchmark tasks and algorithms fairly'. Its five simulated tasks have fixed datasets and a success-rate evaluation; Diffusion Policy calls it 'a large-scale robotic manipulation benchmark'. The library alone would not be a benchmark.
- Built by
- Stanford University, The University of Texas at Austin112
More
Authors: Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, Roberto Martín-Martín. Part of the ARISE Initiative; development began in the Stanford Vision and Learning Lab in late 2018 (README).
Stanford University · 8 of 10 authors (Stanford Vision and Learning Lab)12
The University of Texas at Austin · Soroush Nasiriany and Yuke Zhu1
- Released
- August 2021, at CoRL 2021131415+1
More
arXiv v1 2021-08-06; repository created 2021-08-06; CoRL 2021 (PMLR 164).
- Version
- The library is at 0.5.0 and the datasets are v0.1.1724
More
Library v0.5.0 (released 2025-06-27). Datasets: 'robomimic v0.1' (CoRL 2021), now distributed as robosuite v1.5 versions.
v0.1.0 · Code and paper release dated 2021-08-09 in the README; GitHub release 2021-11-16, 'should be used when trying to reproduce results from the study'217
v0.2.0 (2021-12-17) · Modular observation modalities; MOMART datasets17
v0.3.0 (2023-07-04) · BC-Transformer and IQL; robosuite v1.4 and DeepMind MuJoCo bindings17
v0.4.0 (2025-03-11) · robosuite v1.5 support; datasets moved to Hugging Face17
v0.5.0 (2025-06-27) · Diffusion Policy, multi-dataset training, language-conditioned policies17
Three dataset generations · CoRL 2021 datasets used the mujoco-py 'offline_study' branch of robosuite; v0.3 shipped robosuite 1.4.1 versions; v0.4+ ships robosuite 1.5.1 versions. The docs warn learning results 'may not match exactly'.4
Setup
- Runs in
- Simulation Inferred14
More
Scores in later papers come from the five simulated tasks. The study also ran three real-world versions (Lift, Can, Tool Hang) on a Franka arm; their datasets are released, but real evaluation needs the authors' physical setup.
- Simulator
- robosuite, built on MuJoCo14
More
robosuite on MuJoCo; v1.5.1 recommended since robomimic v0.4. The CoRL 2021 datasets used robosuite's mujoco-py 'offline_study' branch.
Policies act at 20 Hz through an operational space controller.
- Robot
- One arm, Two arms1
- Robot model
- Panda arms in simulation and in the real world; Transport uses two.1
- Setting
- Tabletop1
- Tasks
- 5 simulated and 3 real tasks12
More
8 tasks: 5 simulated (Lift, Can, Square, Transport, Tool Hang) and 3 real (Lift, Can, Tool Hang)
Later papers use the 5 simulated tasks, often only 4 (Tool Hang has no multi-human dataset).
- Training data
- 200 to 300 human demonstrations per task Inferred14
More
3,000 human demonstrations by our arithmetic, plus 5,400 machine-generated trajectories
PH: 200 per task for 5 simulated and 3 real tasks (1,600). MH: 300 per task for 4 tasks (1,200). Paired: 200 (Can). MG: 1,500 (Lift) and 3,900 (Can) from SAC checkpoints.
Proficient-Human (PH) · 200 demos per task from one experienced operator via RoboTurk14
Multi-Human (MH) · 300 demos per task from 6 operators of 3 skill levels (50 each); Lift, Can, Square, Transport14
Machine-Generated (MG) · Rollouts from SAC checkpoints: 1,500 (Lift), 3,900 (Can)4
Paired · 200 Can demos: one success and one failure for each of 100 starts14
Hugging Face size · 26 files, about 6.56 GB (low-dim and raw state files; image versions are generated locally)20
Real datasets · Lift (1.9 GB), Can (5.3 GB), Tool Hang (58 GB) on a Stanford server421
- Changes at test
- Only the start positions change1
More
Object poses are randomized at the start of each episode, within small regions. Objects and scenes stay the same.
Scoring and access
- Scored by
- Success rate1
- Score
- Success rate per task, taken from the best checkpoint saved during training13
More
Success over 50 rollouts, evaluated every few epochs; the study reports the best success over training, averaged over 3 seeds. Results are given per task, per dataset (PH, MH, MG) and per observation type (low-dim or image). There is no standard cross-task average.
Diffusion Policy reports both the best checkpoint and the average of the last 10 checkpoints.
- Error bars
- Sometimes reported Inferred13
More
The study reports mean and standard deviation over 3 seeds. Diffusion Policy gives seed averages without error bars.
- Leaderboard
- None. Scores are only in papers. Inferred2124
More
No leaderboard on the site, docs or README.
- Code licence
- MIT11
More
LICENSE: MIT, copyright 2021 Stanford Vision and Learning Lab.
- Data licence
- Hugging Face dataset card: mit (simulated datasets)20
More
The card says the repository holds 'some of the datasets'. The real-robot datasets on the Stanford server carry no licence statement that we found.
- Asset licence
- MIT, through robosuite, by our reading Inferred22
More
Task assets ship in robosuite, whose LICENSE is MIT (with an Apache-2.0 notice for partial MuJoCo code). No separate asset licence.
Reading the robosuite root licence as covering its bundled meshes is ours; robosuite has no other licence file.
- Access
- Open. The data is on Hugging Face.202123+1
More
Simulated datasets on an ungated Hugging Face repository (moved to robomimic/robomimic_datasets on 2026-02-05; the old amandlek/robomimic id redirects). Real datasets on a Stanford server: Lift and Can links answered HTTP 200 on 2026-10-10; the 58 GB Tool Hang link returned HTTP 503 twice.
Sources 23
- 1robomimic study paper, full text v2 (Tables 1 to 3, Sections 3 and 4, Appendix E)Paper · Sep 2021 · checked 10 Oct 2026
- 2robomimic README (latest updates, description)Repository · Jun 2025 · checked 10 Oct 2026
- 3Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (Tables 1 and 2)Paper · Mar 2023 · checked 10 Oct 2026
- 4robomimic docs: robomimic v0.1 (CoRL 2021) datasets, versions warning, dataset info, reproduction stepsOfficial site · 2025 · checked 10 Oct 2026
- 5Waypoint-Based Imitation Learning for Robotic Manipulation (AWE)Paper · Jul 2023 · checked 10 Oct 2026
- 6Consistency Policy: Accelerated Visuomotor Policies via Consistency DistillationPaper · May 2024 · checked 10 Oct 2026
- 7Diffusion Policy Policy Optimization (DPPO)Paper · Sep 2024 · checked 10 Oct 2026
- 8MimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsPaper · Oct 2023 · checked 10 Oct 2026
- 9RoboCasa v0.2 docs: Policy Learning (official code is a robomimic branch)Repository · Apr 2025 · checked 10 Oct 2026
- 10LIBERO requirements.txt (robomimic==0.2.0)Repository · 10 Dec 2024 · checked 10 Oct 2026
- 11robomimic LICENSE (MIT, 2021 Stanford Vision and Learning Lab)Repository · 2021 · checked 10 Oct 2026
- 12robomimic project siteOfficial site · 2023 · checked 10 Oct 2026
- 13What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (arXiv abstract page)Paper · Aug 2021 · checked 10 Oct 2026
- 14GitHub API: ARISE-Initiative/robomimic (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 15PMLR volume 164 page (Proceedings of the 5th Conference on Robot Learning, 164:1678-1690)Paper · 2022 · checked 10 Oct 2026
- 16robomimic study pageOfficial site · 2021 · checked 10 Oct 2026
- 17robomimic GitHub releases (v0.1.0 to v0.5.0) and release notesRepository · 27 Jun 2025 · checked 10 Oct 2026
- 18robomimic commit history (master)Repository · 9 Aug 2026 · checked 10 Oct 2026
- 19robomimic pull request #294 'update HF_REPO_ID' (datasets moved to a new Hugging Face organisation)Repository · 5 Feb 2026 · checked 10 Oct 2026
- 20Hugging Face dataset robomimic/robomimic_datasets: card (license mit) and file listingRepository · 13 Apr 2025 · checked 10 Oct 2026
- 21robomimic dataset registry (robomimic/__init__.py: horizons, real-data links on Stanford server)Repository · 5 Feb 2026 · checked 10 Oct 2026
- 22robosuite LICENSE (MIT, 2022 Stanford Vision and Learning Lab and UT Robot Perception and Learning Lab)Repository · 2022 · checked 10 Oct 2026
- 23robomimic real-robot datasets on the Stanford download server (Lift, Can, Tool Hang)Repository · 2021 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: Study paper full text (Section 4.7, Appendix E.2); study page; docs; Diffusion Policy, DPPO, MimicGen (real experiments use other tasks); evidence papers PolaRiS (2512.16881), SureSim (2510.04354), 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366), Betting for Sim-to-Real (2604.24018), Robot Policy Evaluation for Sim-to-Real Transfer (2508.11117), Active Real-World Factor-Based Evaluation (2607.14439), Beyond Binary Success (2603.13616): no robomimic mention in full text; web searches. No paired study found.
license_data (real datasets): Docs datasets page, README, Hugging Face card, dataset registry: no licence statement for the Stanford-hosted real datasets.
used_by (count): No tracker found; the docs' 'Projects using robomimic' page only links to Google Scholar.
issues (audits): The 2026 audit (2606.04233) does not cover robomimic; web search for robomimic saturation found AWE's statement only.
Change history
- Created full entry from primary sources (no basic entry existed). Scope decided: in scope as a benchmark of fixed tasks and datasets.