RoboCasa
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
RoboCasa is a kitchen simulator with 100 tasks for a mobile robot arm. Most papers use it to report the average success rate on 24 short tasks.123
What a score here does not tell you Inferred
- How well a policy (the robot's control model) will do on a real robot.No study has compared simulation and real-robot scores for the same policies.
- How well a policy handles long, multi-step chores.The common protocol uses 24 short atomic tasks, each a single basic skill.
- Whether scores from different papers can be compared directly.Papers train on 50 to 3,000 demonstrations per task.
Comparisons with real robots
Tried on real robots, not compared Inferred11718+1
Details
Paper: a Franka on DROID hardware, 3 pick-and-place tasks, 50 real demos each; co-training with all MimicGen data raised average success from 13.6% to 24.4% on seen objects and 2.6% to 9.3% on unseen objects (3 seeds). Simulation and real controllers differ (operational space control at 20 Hz vs 15 Hz). RoboCasa365 reports a similar co-training result. The 2026 audit calls a sim-vs-real ranking test impractical and does not run one. SureSim (2025-10) uses RoboCasa objects in its own ManiSkill3 twin, not RoboCasa tasks.
Our assessment Opinion
Check the training data before you compare RoboCasa averages.
Reasoning
A RoboCasa 24-task average can be compared only with another average from the same training data, the same number of trials, the same checkpoint rule and the same code. Published averages range from 28.8% to 80.6%, and the differences mostly reflect how much and what kind of training data was used. Read the training-data column before the score.
Confidence: high
RoboCasa scores hold up better than LIBERO scores in the 2026 audit. They have not been tested against real robots.
Reasoning
The 2026 audit found no shortcut and a higher share of significant claims than on LIBERO. So a high RoboCasa score is better evidence of skill on these tasks than a high LIBERO score. It is still not evidence of real-world performance. The real-robot work tests whether simulation data helps training. It does not test whether scores predict real results.
Confidence: medium
Scores cover short atomic tasks only. They say little about multi-step chores.
Reasoning
The common protocol uses only the 24 short atomic tasks. The 75 composite tasks, which contain the long, multi-step household chores, are rarely reported, and the original paper's composite results were at most 12%. A RoboCasa score therefore says little about multi-step chores.
Confidence: medium
New work should consider RoboCasa365. If you use the 2024 protocol, fix the code version.
Reasoning
For new work, RoboCasa365 is the maintained successor, and it has an official leaderboard. The 2024 protocol is still useful for comparison with the large body of earlier results, if the code version is fixed.
Confidence: medium
Known problems 6
Results depend on the code version
The same task names exist in v0.1, v0.2, later fixes, forks and the RoboCasa365 code.202122+3
Details
The paper used v0.1, which has no tag. Datasets were recorded with robosuite 1.4.1 while v0.2 runs on robosuite 1.5, and users report version-mismatch warnings. A 2025-02-27 fix on the v0.2 branch changed how object rotations are sampled at reset. NVIDIA's GR00T evaluations run a fork 4 commits ahead of v0.2. Original task names also exist in RoboCasa365 code, where horizons were raised 1.5x in v1.0.1.
Runs are not bit-for-bit repeatable across hardware
Changing only the CPU or GPU changed episode outcomes.1825
Details
With the same StarVLA policy and settings, changing only the CPU made simulator and contact traces diverge at once, with images and actions diverging by step 4 and reward/success by step 19. Changing only the GPU led to reward/success divergence at step 18. Start states are drawn per reset, not per episode index; the maintainers say to pin them with set_ep_meta for paired comparisons.
Training data and trials differ between papers
Papers train on 50 to 3,000 demonstrations per task. Their averages cannot be compared directly.189+7
Details
Papers reporting the 24-task average train on 50 human demos per task (Cosmos Policy, Video Policy, the audit's probe), 300 MimicGen demos (GR00T N1, RLDX-1, Video Policy), 1,000 (World2Act's base) or 3,000 (the paper, DP-VLA). X-WAM pretrains on 56,771 RoboCasa MimicGen episodes before fine-tuning. Trials range from 48 to 150 per task. GR00T N1 and FLARE report the best of the last five checkpoints, each scored on the test scenes; there is no validation split. Z-1 adds RL fine-tuning inside the simulator. The same model can move by 20 points with data alone: GR00T-N1 scores 17.4%, 32.1% and 49.6% with 30, 100 and 300 demos.
The same model gets different numbers in different papers
GR00T N1.5 is reported as 59.7, 64.1 and 65.7 in different papers.101314+5
Details
GR00T N1.5 appears as 64.1 (Cosmos Policy, X-WAM), 65.7 (RLDX-1) and 59.7 (Z-1). Cosmos Policy reports 67.1 for itself; World2Act lists it at 65.7. FLARE reports 70.1; Cosmos Policy lists it at 66.4. pi0 appears as 62.5 with 300 demos (Cosmos Policy, X-WAM, RLDX-1) and 42.4 in Being-H0.7's 50-demo column; GR00T N1.6 as 66.22 (NVIDIA README) and 36.0 (Being-H0.7).
About half of claimed gains can be shown to be statistically significant
16 of 30 claimed improvements can be shown to be statistically significant from the published scores.1827
Details
Of 30 previous-best-to-new comparisons on the 24-task protocol, 16 (53.3%) are provably significant at the 5% level from public scores, 10 are inconclusive and 4 show no improvement. The audit says the share is higher than LIBERO's 19.8% partly because RoboCasa has fewer reported results, so scores sit further apart.
The asset licence may not match the terms of the original sources
RoboCasa licenses its assets under CC BY 4.0. It does not publish the licence terms of the original objects. Inferred52829+1
Details
RoboCasa licenses all assets under CC BY 4.0, but its Objaverse objects carry individual licences, and Objaverse includes non-commercial and share-alike objects. RoboCasa does not list per-object licences. ManiSkill's copy of the scenes is labelled MIT.
Details
About
- What it is
- Kitchen simulator with fixed tasks Inferred1
More
The paper calls RoboCasa a simulation framework and defines 100 tasks 'for systematic evaluation'. Later papers use its 24 atomic tasks as a benchmark.
- Built by
- The University of Texas at Austin, NVIDIA Research1
- Released
- June 2024, at RSS 202431632
More
arXiv v1 2024-06-04 (only version); repository created 2024-05-11; RSS 2024.
The RSS proceedings title says 'Household Tasks' where arXiv says 'Everyday Tasks'.
- Version
- v0.2 (October 2024)53322+1
More
v0.2 (2024-10-31): robosuite v1.5 backend. Tagged on 2025-12-18 and published as a GitHub release on 2026-02-17. The 2024 paper used v0.1, which has no tag.
v0.1 (2024-06) · No tag. The maintainer points to commit 0f604b25 (2024-10-28) and the robocasa_v0.1 branch of robosuite.20
v0.2 branch changes · Bug fix on 2025-02-27 changes how object rotations are sampled at reset (quaternion order in the placement sampler); default renderer changed 2025-03-18.22
NVIDIA fork · Isaac-GR00T evaluates 'RoboCasa Kitchen' with squarefk/robocasa, 4 commits ahead of v0.2 (gymnasium wrapper; reward returns 1 on success), last commit 2025-11-06.2324
RoboCasa365 v1.0 (2026-02-18) · Successor on the same repository's main branch; v1.0.1 (2026-05-12) raised task horizons 1.5x. See facts.successor.233
- Last update
- March 2025. The repository then moved on to RoboCasa365.2233
More
Last code change on the v0.2 branch 2025-03-18 (default renderer); docs fix 2025-04-23; README edit 2025-12-18. Main moved to RoboCasa365 on 2026-02-18.
- Status
- Replaced by RoboCasa365. Its 24-task protocol is still widely used. Inferred2333
More
Superseded by RoboCasa365 in the same repository. Its 24-task protocol is still widely reported.
The main README calls RoboCasa365 'the latest iteration' of RoboCasa. The audit tracker counts 12 papers in March 2026 and 11 in April 2026 reporting RoboCasa results (all protocols).
Setup
- Runs in
- Simulation1
More
Real-robot experiments in the paper only test co-training with simulation data.
- Simulator
- robosuite, built on MuJoCo1520
More
robosuite on MuJoCo (v0.1: robosuite robocasa_v0.1 branch; v0.2: robosuite v1.5). Optional NVIDIA Omniverse rendering.
Paper: about 25.2 steps per second with rendering, roughly real time.
- Robot
- Arm on wheels1
- Robot model
- All experiments use a Franka Panda on an Omron mobile base. The framework also supports other mobile manipulators, humanoids and quadrupeds with arms.1
- Setting
- Kitchen1
- Tasks
- 100 tasks, of which 24 are in common use1810
More
100 tasks: 25 atomic and 75 composite (composite tasks proposed with LLM help). Common protocol: 24 atomic tasks (navigation excluded).
30 tasks with released datasets · 25 atomic and 5 composite tasks have human datasets; 24 atomic tasks have MimicGen datasets (counted by us from the v0.2 registry)3435
Composite results · The paper reports 5 composite tasks, single-task policies on 50 human demos: 0 to 2.0% from scratch, 0 to 12.0% with fine-tuning1
- Scenes
- 120 kitchens, of which 5 are used for testing13617+1
More
120 kitchen scenes: 10 floor plans x 12 styles. Evaluation uses 5 fixed scenes.
The v0.2 code has 10 layout and 12 style files besides a 'playground' test file (our count). CONFLICT: the RoboCasa365 paper describes the original as having 100 scenes. The audit lists the five evaluation layout/style pairs as (1,1), (2,2), (4,4), (6,9), (7,10).
- Training data
- 50 human demonstrations per task, and more than 100,000 generated ones Inferred13534
More
1,500 human demonstrations by our arithmetic (50 per task for 25 atomic and 5 composite tasks) plus 100K+ MimicGen trajectories.
Human · 1,250 for the 25 atomic tasks, collected by 4 operators with a SpaceMouse in random kitchens1
MimicGen · 72,000 trajectories (3,000 per task, 24 atomic tasks) with Objaverse objects, used in the paper, plus 28K with AI-generated objects1
Quality caveat · The paper says many generated trajectories show jerky motions and collisions1
- Changes at test
- New objects and new kitchen styles Inferred125
More
Evaluation uses only unseen object instances; two of five evaluation kitchens have styles never seen in training. Object placements are sampled at each reset.
Object-instance and unseen styles are stated in the paper; mapping unseen styles to 'visual' and placement sampling to 'object-pose' is ours. Kitchen layouts are not held out: training demos come from random layouts.
Scoring and access
- Scored by
- Success rate1
- Score
- Success averaged over 24 tasks1810+1
More
Average success over the 24 atomic tasks, each tested on unseen objects in five fixed kitchens. Training data, trial count and checkpoint choice are not fixed.
The official policy-learning code is a robomimic branch with BC-Transformer. The paper's main result trains one multi-task BC-Transformer per dataset size and reports per-task and overall success.
- Trials
- 50 per task in the paper. Other papers use 48 to 150.189+5
More
Paper · 50 trials per task across 5 fixed scenes1
GR00T N1 · 100 trials; maximum of the last 5 checkpoints8
FLARE · 50 episodes per task; maximum over the final 5 checkpoints9
Cosmos Policy · 50 trials per task x 3 seeds (3,600 trials)10
World2Act · 50 trials per task x 5 seeds12
X-WAM · 100 episodes per task13
Z-1 · 48 rollouts per task (SFT), 64 (RL)16
2026 audit probe · 50 rollouts per task, 1,200 in total18
- Who runs it
- Each team tests its own model Inferred538
More
No organiser runs the 2024 protocol. The official leaderboard covers RoboCasa365 only.
- Error bars
- Not reported Inferred1810+2
More
None of the papers we read gives a confidence interval or standard deviation for the 24-task average. The original paper gives mean and standard deviation only for its real-robot co-training results.
- Leaderboard
- None. Scores are only in papers. Inferred53839
More
No official leaderboard for the 2024 protocol. A community page (GINIGEN-AI, created 2026-06-29, not updated since) lists six matched-protocol entries copied from the RLDX-1 report.
- Code licence
- MIT456
More
LICENSE: MIT, copyright 2024 the RoboCasa Team, with a notice that partial MuJoCo code is under Apache-2.0. The GitHub API reports NOASSERTION because of the extra notice.
- Data licence
- CC-BY-4.05
More
README licence section: 'Assets and Datasets: CC BY 4.0'.
- Asset licence
- CC BY 4.0. The terms of the original object sources were not checked.5128+1
More
README: CC BY 4.0. Upstream terms of individual objects were not checked (items).
Objaverse objects · Objaverse objects are individually licensed; its card lists 25K CC-BY-NC, 52K CC-BY-NC-SA and 16K CC-BY-SA objects among 800K+. RoboCasa does not say which licences its 914 to 917 Objaverse objects carry.2829
Fixtures and AI-generated objects · Fixtures come from 'online 3D model repositories' and most objects from Luma.ai; no per-source terms given1
ManiSkill copy · ManiSkill's copy of RoboCasa scenes on Hugging Face is labelled MIT30
Sources 40
- 1RoboCasa paper, full text v1 (Sections III to V, Appendix VIII and IX, Figures 7, 10, 11, 13)Paper · Jun 2024 · checked 10 Oct 2026
- 2RoboCasa README on main (RoboCasa365 updates, licence, citations)Repository · 25 Sep 2026 · checked 10 Oct 2026
- 3Audit artifacts: RoboCasa citation tracker CSV (snapshot 2026-05-21)Repository · May 2026 · checked 10 Oct 2026
- 4RoboCasa LICENSE at tag v0.2 (MIT, with Apache-2.0 notice for partial MuJoCo code)Repository · 2024 · checked 10 Oct 2026
- 5RoboCasa README at tag v0.2 (latest updates, licence section)Repository · 18 Dec 2025 · checked 10 Oct 2026
- 6GitHub API: robocasa/robocasa (stars, forks, created, licence field)Index · 10 Oct 2026 · checked 10 Oct 2026
- 7A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM (DP-VLA, Table 1)Paper · Oct 2024 · checked 10 Oct 2026
- 8GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (Table 4, evaluation protocol)Paper · Mar 2025 · checked 10 Oct 2026
- 9FLARE: Robot Learning with Implicit World Modeling (Table 1)Paper · May 2025 · checked 10 Oct 2026
- 10Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning (Table 2)Paper · Jan 2026 · checked 10 Oct 2026
- 11Video Generators are Robot Policies (Table 1)Paper · Aug 2025 · checked 10 Oct 2026
- 12World2Act: Latent Action Post-Training from World Model Dynamics (Table 1)Paper · Mar 2026 · checked 10 Oct 2026
- 13Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising (X-WAM, Table 1, Appendix B)Paper · Apr 2026 · checked 10 Oct 2026
- 14RLDX-1 Technical Report (Table 1b, benchmark descriptions)Paper · May 2026 · checked 10 Oct 2026
- 15NVIDIA Isaac-GR00T: RoboCasa evaluation example and checkpoint results (GR00T N1.6, N1.7)Repository · 26 May 2026 · checked 10 Oct 2026
- 16Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models (Table 1, Appendix C.5)Paper · Jun 2026 · checked 10 Oct 2026
- 17RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist RobotsPaper · Mar 2026 · checked 10 Oct 2026
- 18What Are We Actually Benchmarking in Robot Manipulation? (Table 1, Figure 3, Appendix A.1, B, D)Paper · Jun 2026 · checked 10 Oct 2026
- 19Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators (SureSim; uses RoboCasa objects)Paper · Oct 2025 · checked 10 Oct 2026
- 20RoboCasa issue #143 'Robocasa v0.1 where can I get it from?' (maintainer reply)Repository · 12 May 2025 · checked 10 Oct 2026
- 21RoboCasa issue #144 'Potential Inconsistency in Dataset and Policy Learning Repo'Repository · 13 May 2025 · checked 10 Oct 2026
- 22RoboCasa v0.2 branch commit history (incl. placement-sampler fix f202a7eb, 2025-02-27)Repository · 18 Dec 2025 · checked 10 Oct 2026
- 23NVIDIA Isaac-GR00T .gitmodules (robocasa submodule from squarefk/robocasa)Repository · 2026 · checked 10 Oct 2026
- 24GitHub compare robocasa v0.2 ... squarefk/robocasa@d89d481c (4 commits ahead, 1 behind)Repository · 6 Nov 2025 · checked 10 Oct 2026
- 25RoboCasa issue #218 'Is there a supported way to hold the initial-state sequence fixed across two runs?'Repository · 25 Aug 2026 · checked 10 Oct 2026
- 26Being-H0.7: A Latent World-Action Model from Egocentric Videos (Table 1)Paper · Apr 2026 · checked 10 Oct 2026
- 27Audit artifacts: statistical significance pie counts CSVRepository · May 2026 · checked 10 Oct 2026
- 28Objaverse dataset card (licence breakdown of individual objects)Repository · 2023 · checked 10 Oct 2026
- 29RoboCasa v0.2 docs: Objects (per-category Objaverse and AI-generated counts)Repository · Oct 2024 · checked 10 Oct 2026
- 30Hugging Face dataset card haosulab/RoboCasa (ManiSkill copy of RoboCasa scenes, labelled mit)Repository · 2024 · checked 10 Oct 2026
- 31RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots (arXiv abstract page)Paper · Jun 2024 · checked 10 Oct 2026
- 32Robotics: Science and Systems XX, paper 50 (title 'RoboCasa: Large-Scale Simulation of Household Tasks for Generalist Robots')Paper · Jul 2024 · checked 10 Oct 2026
- 33RoboCasa GitHub releases (v0.2 'Original RoboCasa Release', v1.0 'RoboCasa365 Release')Repository · 18 Feb 2026 · checked 10 Oct 2026
- 34RoboCasa v0.2 dataset registry (25 single-stage and 5 multi-stage task datasets; UT Box links)Repository · Oct 2024 · checked 10 Oct 2026
- 35RoboCasa v0.2 docs: Downloading DatasetsRepository · Oct 2024 · checked 10 Oct 2026
- 36RoboCasa v0.2 kitchen layout and style filesRepository · Oct 2024 · checked 10 Oct 2026
- 37RoboCasa v0.2 docs: Policy Learning (robomimic 'robocasa' branch, BC-Transformer)Repository · Apr 2025 · checked 10 Oct 2026
- 38RoboCasa365 leaderboard pageLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 39RoboCasa Kitchen Leaderboard (community Hugging Face Space by GINIGEN-AI)Secondary · 29 Jun 2026 · checked 10 Oct 2026
- 40RoboCasa v0.2 asset download script (UT Box links)Repository · Oct 2024 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: RoboCasa paper full text (Section V-C); RoboCasa365 paper; 2026 audit (2606.04233; no sim-vs-real test); PolaRiS (2512.16881; related-work mention only); 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366; related work only); SureSim (2510.04354; uses RoboCasa objects in its own twin); Betting for Sim-to-Real (2604.24018), Robot Policy Evaluation for Sim-to-Real Transfer (2508.11117), Active Real-World Factor-Based Evaluation (2607.14439), Beyond Binary Success (2603.13616): no RoboCasa mention; audit tracker notes; web searches. No paired study found.
top_score (anything above 79.2% without RL): Audit tracker (33 protocol-comparable rows), Cosmos Policy and X-WAM comparison tables, NVIDIA Isaac-GR00T README, community leaderboard, web search on 2026-10-10. Tracker rows above 79.2 use custom or partial task sets.
license_assets (per-object terms): README, paper, v0.2 docs objects page, Hugging Face asset card (no text). No per-object licence list found.
uncertainty_reported: RoboCasa paper, GR00T N1, FLARE, Video Policy, Cosmos Policy, World2Act, X-WAM, RLDX-1, Z-1, DP-VLA result tables.
Change history
- Created full entry from primary sources, starting from the checked basic entry and core-sim-a notes. Verified top scores at their sources (12 results), added the 2026 audit's tracker counts and significance breakdown, protocol and version issues, successor relation to RoboCasa365. Corrected: latest code change on v0.2 (2025-03, not only the 2024-10 README date); uncertainty reporting; access (Box links alive).