RoboCasa365
Kitchen simulation benchmark with 365 tasks in 2,500 scenes; its leaderboard ranks policies on 50 target tasks.1
Comparisons with real robots
Tried on real robots, not compared1
Details
Training-data evidence only: GR00T N1.5 mid-trained on 150 sim tasks then co-fine-tuned with 140 real demos reached 79.8% average real success vs 61.8% real-only (4 kitchen tasks, 20 trials each, DROID Panda; by the authors). No study pairing RoboCasa365 sim scores with real scores for several policies found.
Known problems 2
Top entry's submission says it is 'a full fine-tune of xiaomi-robotics-1' with paper_link
Our reading of the submission JSON against the README rules; the organisers may have accepted it under the 'new recipe' clause. Inferred4
V1.0.1 lengthened horizons 1.5x; gr00t n1.5 re-evaluated (leaderboard 23.9 vs docs table 2
Docs table: https://raw.githubusercontent.com/robocasa/robocasa/main/docs/benchmarking/multitask_learning.md.5
Details
About
- What it is
- Benchmark Inferred6
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- The University of Texas at Austin; NVIDIA Research1
- Released
- 2026-027
More
GitHub release v1.0 (RoboCasa365) 2026-02-18; arXiv v1 2026-03-04.
- Version
- 1.0.18
More
setup.py and robocasa/__init__.py; no GitHub release object for 1.0.1.
- Last update
- v1.0.1 (2026-05-12): all task horizons increased 1.5x; 2026-07-07: per-frame subtask annotations for target composite datasets; leaderboard updated 2026-10-109
More
README 'Updates'; leaderboard page header.
- Status
- Active Inferred5
More
Leaderboard updated 2026-10-10; commits 2026-09.
Setup
- Runs in
- Simulation1
- Robot
- Arm on wheels1
More
Franka Panda with Omron mobile base for data collection; framework 'in principle' supports other mobile manipulators and humanoids.
- Setting
- Kitchen1
- Size
- 365 tasks (65 atomic, 300 composite); 2,500 pretraining kitchens + 10 target kitchens; 3,200+ objects; 30k human pretraining demos; 10k MimicGen demos per task for 60 atomic tasks; 25k target demos1
Scoring and access
- Scored by
- Success rate5
More
Weighting inferred: Axiom-0 (86.5*18 + 59.4*16 + 35.8*16)/50 = 61.6, matching the page.
- Trials
- Conflict: 30 trials per task (paper App. G.2) vs 50 sampled scenarios per task (docs benchmarking overview)10
More
Paper: https://arxiv.org/pdf/2603.04356.
- Who runs it
- both: self-submitted JSON via pull request; RoboCasa team verifies (up to 10 days); closed models must give the team private access to checkpoint and eval code11
- Leaderboard
- Official leaderboard5
More
Published 2026-04-06 (README). 16 models shown 2026-10-10; submissions folder holds 19 JSON files (our count).
- Code licence
- MIT3
More
Copyright (c) 2026 the RoboCasa Team.
- Data licence
- CC BY 4.0 (assets and datasets)9
More
Hugging Face robocasa/robocasa-assets card: cc-by-4.0.
- Commercial use
- Allowed Inferred9
More
Upstream Objaverse object terms not checked. Not legal advice.
Sources 11
- 1RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist RobotsPaper · Mar 2026 · checked 10 Oct 2026
- 2Semantic Scholar API recordIndex · checked 10 Oct 2026
- 3robocasa/robocasa on GitHub (main)Repository · checked 10 Oct 2026
- 4robocasa-benchmark/leaderboard on GitHub (file Axiom-0_2026_10_01.json)Leaderboard · checked 10 Oct 2026
- 5RoboCasa LeaderboardLeaderboard · checked 10 Oct 2026
- 6RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist RobotsPaper · Mar 2026 · checked 10 Oct 2026
- 7robocasa/robocasa on GitHub (releases)Repository · checked 10 Oct 2026
- 8robocasa/robocasa on GitHub (file setup.py)Repository · checked 10 Oct 2026
- 9robocasa/robocasa on GitHub (file README.md)Repository · checked 10 Oct 2026
- 10robocasa/robocasa on GitHub (file benchmarking_overview.md)Repository · checked 10 Oct 2026
- 11robocasa-benchmark/leaderboard on GitHub (repository)Leaderboard · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).