| 1X World Model ChallengeCompetition to predict future first-person frames of 1X's EVE robot from its logs and actions. |
World models |
Recorded data |
Not checked |
Humanoid |
2024 |
| 1X World Model eval1X's internal video world model that predicts humanoid task success to rank policy checkpoints before real trials. |
World models |
AI simulator |
Not checked |
Humanoid |
2025 |
| AgiBot WorldAgiBot World is a training dataset recorded with about 100 real AgiBot robots. It is used to pretrain robot policies (the models that control robots), and it has no public test of its own. |
Handling objects |
Real robots |
Real robots |
Humanoid, Two arms, Robot hand |
2024 |
| AgiBot World 2026AgiBot's 2026 real-robot dataset on its G2 robot, released in themed phases; no evaluation protocol. |
Handling objects |
Real robots |
Real robots |
Two arms, Robot hand |
2026 |
| AgiBot World Challenge · R2AAgiBot manipulation contest (2025, 2026): online simulation ranking on 10 tasks, then finalists run on real robots. |
Handling objects |
Sim + real |
Real robots |
Humanoid, Arm on wheels |
2025 |
| AgiBot World Challenge · WMAgiBot contest track (2025, 2026): predict robot head-camera video from actions; scored against held-out real video. |
World models |
Recorded data |
Not checked |
Humanoid |
2025 |
| AI2-THORUnity-based simulator of interactive indoor rooms and houses used to train and test navigation and manipulation agents. |
Navigation |
Simulation |
Not checked |
Wheels only, Arm on wheels, Game character |
2017 |
| AI2-THOR RearrangementSimulated task where an agent must restore moved objects in a room to their earlier positions and states. |
Household tasks |
Simulation |
Not checked |
Arm on wheels |
2021 |
| ALFREDALFRED is a benchmark in which an AI agent follows written instructions to do household tasks in simulated rooms. Agents are scored on test rooms they did not see in training. |
Household tasks |
Simulation |
Not checked |
Game character |
2019 |
| ASIMOVASIMOV is a set of question sets from Google DeepMind. It tests whether AI models recognise physical danger the way human raters do. |
Safety |
Recorded data |
Not checked |
No body |
2025 |
| ASIMOV-AgenticTests whether a robot's planning model refuses unsafe tasks, stops near people and asks for help when unsure. |
Safety |
Recorded data |
Not checked |
No body |
2026 |
| AutoEvalAutoEval is a system that tests submitted robot policies (the models that control robots) on real robot arms without a human operator. Its public robot stations, called cells, are now offline. |
Handling objects |
Real robots |
Real robots |
One arm |
2025 |
| BEHAVIOR ChallengeAnnual simulation competition on 50 (2025) or 100 (2026) BEHAVIOR-1K household tasks with a bimanual mobile robot. |
Household tasks |
Simulation |
Not checked |
Arm on wheels, Two arms |
2025 |
| BEHAVIOR-1KBEHAVIOR-1K is a set of 1,000 household activities in simulation. Its yearly challenge scores a two-armed mobile robot on a subset of these activities. |
Household tasks |
Simulation |
Not checked |
Arm on wheels |
2022 |
| BridgeData V2BridgeData V2 is a set of 60,096 robot trajectories (recorded robot runs) on a low-cost arm, mostly in toy kitchens. It is used as training data, and its robot setup is reused for tests. |
Handling objects |
Real robots |
Real robots |
One arm |
2023 |
| CALVINCALVIN is a simulated benchmark in which a robot arm at a desk follows chains of five spoken-style instructions. It tests how many of the five tasks a policy (the robot's control model) completes in a row. |
Handling objects |
Simulation |
Not checked |
One arm |
2021 |
| CausalVQAVideo question set on cause and effect in real Ego-Exo4D activity clips: counterfactual, anticipation, planning. |
Reasoning |
Recorded data |
Not checked |
No body |
2025 |
| CHORESTen-task simulated household benchmark where a Stretch robot finds, picks up and fetches objects in generated houses. |
Household tasks |
Simulation |
Not checked |
Arm on wheels |
2023 |
| Cosmos-HumanEvalNVIDIA human-rating protocol: annotators answer yes/no physics and fidelity questions about generated videos. |
World models |
Recorded data |
Not checked |
No body |
2026 |
| DialFREDALFRED extension where the agent may ask a human questions about locations, appearance, or direction to finish tasks. |
Working with people |
Simulation |
Not checked |
Game character |
2022 |
| DROIDDROID is a dataset of 76k real robot demonstrations recorded in 564 scenes on one shared robot setup. It has no fixed test. Labs use the same setup to test policies (the models that control robots). |
Handling objects |
Real robots |
Real robots |
One arm |
2024 |
| Ego-Exo4D ProficiencyVideo benchmark: estimate a person's skill level, and spot good or weak moments, from first- and third-person video. |
Reasoning |
Recorded data |
Not checked |
No body |
2023 |
| Ego4D3,670 hours of first-person human video with a benchmark suite on memory, hand-object interaction, social cues and forecasting. |
Reasoning |
Recorded data |
Not checked |
No body |
2021 |
| EmbodiedBenchEmbodiedBench is a set of 1,128 simulated tasks in which a multimodal AI model (one that takes in both images and text) plans a robot's actions. It reports the share of tasks the model completes in each of its four environments. |
Household tasks |
Simulation |
Not checked |
Game character, Arm on wheels, Wheels only, One arm |
2025 |
| EmbodiedGovBenchScores embodied agent systems on governance: permission limits, recovery, upgrades, human override and audit trails, in AI2-THOR scenarios. |
Safety |
Simulation |
Not checked |
— |
2026 |
| ERQAERQA is a set of 400 multiple-choice questions about images from robots and first-person video. It scores how well vision-language models (AI models that read images and text) answer them, and no robot moves. |
Reasoning |
Recorded data |
Not checked |
No body |
2025 |
| Eval-ActionsReal-robot episodes with expert quality grades, used to train and test models that judge how well manipulation was executed. |
Handling objects |
Recorded data |
Not checked |
One arm, Two arms |
2026 |
| EWMBenchEWMBench is a benchmark for world models, which here are AI models that generate videos of a robot doing a task. It scores how closely those videos match real recordings of the same task. |
World models |
Recorded data |
Not checked |
Two arms, Humanoid |
2025 |
| Gemini Robotics evalsGoogle DeepMind's in-house real-robot tests of its Gemini Robotics models; task suites are described but not released. |
Handling objects |
Real robots |
Real robots |
Two arms, Humanoid, Robot hand, Many types |
2025 |
| Genie SimAgiBot's simulation benchmark scoring robot policies on instruction, spatial, robustness and manipulation boards, with a VLM judge. |
Handling objects |
Simulation |
Checked |
Humanoid, Two arms |
2025 |
| GenManipIsaac Sim tabletop platform that generates instruction-following manipulation tasks with an LLM; its benchmark has 200 scenarios. |
Handling objects |
Simulation |
Not checked |
One arm, Two arms |
2025 |
| GRUtopiaIsaac Sim platform of large interactive scenes with LLM-driven characters; its benchmark tests legged robots navigating and fetching objects. |
Navigation |
Simulation |
Not checked |
Humanoid, Legged, Arm on wheels |
2024 |
| Habitat 3.0Habitat 3.0 is a simulator release with two tasks in which a simulated Spot robot works with simulated people. The robot either finds and follows a person or tidies a home together with one. |
Working with people |
Simulation |
Not checked |
Arm on wheels |
2023 |
| Habitat Navigation ChallengeThe Habitat Navigation Challenge was a yearly contest from 2019 to 2023. A simulated robot had to reach a given point or object inside scanned buildings it had not seen before. |
Navigation |
Simulation |
Checked |
Wheels only |
2019 |
| HM3D-OVONSimulation benchmark for finding objects named by free-form text, across 379 categories in 3D-scanned homes. |
Navigation |
Simulation |
Not checked |
Wheels only |
2024 |
| HumanoidBenchHumanoidBench is a set of 27 simulated tasks for a humanoid robot with two hands. It tests reinforcement learning methods (which learn by trial and error from a reward) and scores them by the reward they collect. |
Walking and balance |
Simulation |
Not checked |
Humanoid, Robot hand |
2024 |
| HumanTrackerScores humanoid motion trackers in MuJoCo on 153 hours of mocap by success, pose error and a learned preference metric. |
Walking and balance |
Simulation |
Not checked |
Humanoid |
2026 |
| HumEnvSimulated SMPL humanoid environment with reward, goal-reaching and motion-tracking test suites for whole-body control policies. |
Walking and balance |
Simulation |
Not checked |
Game character |
2024 |
| IntPhys 2Tests whether video models can tell physically possible from impossible events in Unreal Engine videos. |
Reasoning |
Recorded data |
Not checked |
No body |
2025 |
| Isaac Lab-ArenaNVIDIA framework for composing simulated robot tasks and running large-scale policy evaluations on Isaac Lab. |
Handling objects |
Simulation |
Not checked |
One arm, Humanoid, Robot hand |
2026 |
| LIBERO130 simulated tasks for one robot arm. Many papers use it to report results for vision-language-action (VLA) models. |
Handling objects |
Simulation |
Checked |
One arm |
2023 |
| LocoMuJoCoImitation-learning benchmark for locomotion: humanoid, quadruped and human-body models imitate motion-capture data in MuJoCo. |
Walking and balance |
Simulation |
Not checked |
Humanoid, Legged |
2023 |
| ManipulaTHORSimulated task where a mobile arm must fetch an object and carry it to a target point in a kitchen. |
Household tasks |
Simulation |
Not checked |
Arm on wheels |
2021 |
| ManiSkillSimulated benchmark of 4 tasks on 162 articulated objects, testing whether manipulation skills transfer to unseen object instances. |
Household tasks |
Simulation |
Not checked |
Arm on wheels, Two arms |
2021 |
| ManiSkill2Simulated benchmark of 20 manipulation task families, rigid and soft body, with 2,000+ objects and 4M+ demo frames. |
Handling objects |
Simulation |
Not checked |
One arm, Arm on wheels, Two arms |
2022 |
| ManiSkill3ManiSkill3 is an open-source robot simulator that runs many copies of a task in parallel on one GPU. Its documentation lists 51 tasks, and it has no single benchmark score. |
Handling objects |
Simulation |
Checked |
One arm, Arm on wheels, Humanoid, Legged, Robot hand, Two arms |
2024 |
| Meta Motivo studyFifty raters compared videos of two simulated-humanoid policies on 95 tasks for task success and human-likeness. |
Moving like people |
Simulation |
Not checked |
Game character |
2024 |
| Meta-WorldMeta-World is a set of 50 simulated tabletop tasks for one Sawyer robot arm. It is used to test multi-task and meta reinforcement learning (learning to adapt quickly to new tasks), and recent vision-language-action (VLA) papers also report results on it. |
Handling objects |
Simulation |
Not checked |
One arm |
2019 |
| Motion Turing TestPeople rate how human-like humanoid robot motions look (0-5), using pose-only replays of 1,000 human and robot clips. |
Moving like people |
Recorded data |
Not checked |
Humanoid |
2026 |
| MultiONSimulation benchmark where an agent must find several target objects in a set order in 3D homes. |
Navigation |
Simulation |
Not checked |
Wheels only |
2020 |
| MVPBenchVideo question set on physical understanding; each question comes as a near-identical video pair with opposite answers. |
Reasoning |
Recorded data |
Not checked |
No body |
2025 |
| Open X-EmbodimentOpen X-Embodiment (OXE) pools robot data from dozens of labs into one format, and model builders use it as training data. It has no test of its own. |
Handling objects |
Real robots |
Real robots |
Many types |
2023 |
| PAI-BenchTests video generators and video-language models on real-world physical-AI clips: driving, robots, industry, people. |
World models |
Recorded data |
Not checked |
No body |
2025 |
| PARTNRPARTNR is a set of 100,000 simulated household tasks in which a robot's planner (the model that decides its next steps) works with a person. It is mainly used to test task planning and coordination by large language models (LLMs). |
Working with people |
Simulation |
Not checked |
Arm on wheels |
2024 |
| PolaRiSPolaRiS builds simulated test scenes from short video scans of real scenes, to score robot policies (the robots' control models) for the DROID robot setup. Its authors report that the scores track real-robot results after each policy is briefly fine-tuned with some simulated data (co-training). |
Handling objects |
Simulation |
Checked |
One arm |
2025 |
| ProcTHORGenerator of procedurally built interactive houses in AI2-THOR, with a 10,000-house training set and a 10-house test set. |
Navigation |
Simulation |
Not checked |
Wheels only, Arm on wheels |
2022 |
| RBenchScores video generators on robot task videos across five task types and four robot body types. |
World models |
Recorded data |
Not checked |
One arm, Two arms, Humanoid, Legged |
2026 |
| RLBenchRLBench is a set of 100 simulated tasks for one Franka Panda robot arm. Papers on robot manipulation policies (the robot's control models) use it, and most report results on an 18-task subset. |
Handling objects |
Simulation |
Not checked |
One arm |
2019 |
| RoboArenaRoboArena is a service that tests robot policies (the robots' control models) on real DROID robot arms. Volunteers run blind A/B tests, in which two unnamed policies try the same task, and their preferences are combined into a rating. |
Handling objects |
Real robots |
Real robots |
One arm |
2025 |
| RoboCasaRoboCasa is a kitchen simulator with 100 tasks for a mobile robot arm. Most papers use it to report the average success rate on 24 short tasks. |
Handling objects |
Simulation |
Not checked |
Arm on wheels |
2024 |
| RoboCasa365Kitchen simulation benchmark with 365 tasks in 2,500 scenes; its leaderboard ranks policies on 50 target tasks. |
Household tasks |
Simulation |
Not checked |
Arm on wheels |
2026 |
| RoboChallengeRoboChallenge tests robot policies (the models that control a robot) on real robots in Dexmal's lab. Users control the robots over the internet, and lab staff set up and score each test on 30 tabletop tasks. |
Handling objects |
Real robots |
Real robots |
Two arms, One arm |
2025 |
| RoboDojoBimanual manipulation benchmark with 42 simulated and 18 real-robot tasks and an organiser-run leaderboard. |
Handling objects |
Sim + real |
Real robots |
Two arms |
2026 |
| RoboLab-120Simulation benchmark of 120 tabletop manipulation tasks for generalist robot policies trained on real-world data. |
Handling objects |
Simulation |
Checked |
One arm |
2026 |
| robomimicrobomimic is a set of five simulated robot-arm tasks with human demonstrations, used to test imitation-learning methods (methods that learn by copying demonstrations). The easier tasks are now near 100% success. |
Handling objects |
Simulation |
Not checked |
One arm, Two arms |
2021 |
| RoboMINDRoboMIND is a dataset of about 107,000 demonstrations on four robot types, used to train robot control models (policies). About 28% of the demonstrations come from simulation, and its paper tests policies on real robots. |
Handling objects |
Real robots |
Real robots |
One arm, Two arms, Humanoid, Robot hand |
2024 |
| RoboMIND 2.0Dual-arm and mobile manipulation demonstrations on six robot types, with tactile data and a small Isaac Sim benchmark. |
Handling objects |
Real robots |
Real robots |
Two arms, Arm on wheels, Humanoid, Robot hand |
2025 |
| RoboTHOR ObjectNavObject-goal navigation benchmark in 89 simulated apartments, 14 of them rebuilt physically for robot testing. |
Navigation |
Simulation |
Not checked |
Wheels only |
2020 |
| RoboTwin 2.0RoboTwin 2.0 is a simulated benchmark and data generator for two-armed robots. It tests robot policies (the models that control a robot) on 50 tabletop tasks in clean and randomized scenes, and it has an official leaderboard. |
Handling objects |
Simulation |
Not checked |
Two arms |
2025 |
| RoboWorldRuns robot policies inside a learned video world model trained on DROID; a VLM judge scores task progress. |
Handling objects |
AI simulator |
Checked |
One arm |
2026 |
| SIMA 2 evaluationGoogle DeepMind's internal test of a Gemini-based game agent: task success in 3D video games, compared with human players. |
Navigation |
Simulation |
Not checked |
Game character |
— |
| SimplerEnvSimplerEnv is a set of simulated copies of two real robot test setups, used to score robot policies (the models that control a robot) in simulation. Its scores agreed with real-robot results for policies from 2023 and 2024, and later checks with newer policies found weaker agreement. |
Handling objects |
Simulation |
Checked |
One arm, Arm on wheels |
2024 |
| TEAChDataset and benchmarks of human-human dialogues for completing household tasks together in the AI2-THOR simulator. |
Working with people |
Simulation |
Not checked |
Game character |
2021 |
| UMI-BenchReal-robot benchmark of 10 tabletop tasks for policies trained on handheld UMI-style gripper data. |
Handling objects |
Real robots |
Real robots |
One arm, Two arms |
2026 |
| VLABenchSimulated single-arm benchmark of 100 language-driven manipulation task types testing common sense, semantics and long-horizon planning. |
Handling objects |
Simulation |
Not checked |
One arm |
2024 |
| VLN-CEVLN-CE is a simulated benchmark in which an agent follows written route instructions through 3D scans of real buildings, using small, robot-like moves. It is widely used to test language-guided navigation models. |
Navigation |
Simulation |
Not checked |
Wheels only |
2020 |
| VSI-BenchVSI-Bench is a set of 5,130 questions about videos of real rooms, used to test how well multimodal AI models (models that take in video and text) understand space. Models only answer questions, and no robot moves. |
Reasoning |
Recorded data |
Not checked |
No body |
2024 |
| WAGIBenchTests whether vision-language models can guess a smart-glasses wearer's goal from video, audio and phone context. |
Reasoning |
Recorded data |
Not checked |
No body |
2025 |
| YD/T 6770-2026Chinese telecom-industry standard defining how to benchmark an embodied AI system in simulation and on real hardware. |
— |
Sim + real |
Real robots |
— |
Issu |