EmbodiedBench
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
EmbodiedBench is a set of 1,128 simulated tasks in which a multimodal AI model (one that takes in both images and text) plans a robot's actions. It reports the share of tasks the model completes in each of its four environments.12
What a score here does not tell you Inferred
- How well a model will do on a real robot.EmbodiedBench runs only in simulation. No study has compared its scores with real-robot results.
- Whether a model can control a robot's fine movements.The model picks from ready-made skills or gives arm poses that a motion planner carries out.
- How good the model is on its own.Scores change with the prompts, the examples and any memory add-ons used around the model.
Comparisons with real robots
Details
The paper's Limitations section says evaluation is solely in simulation. Vlaser (ICLR 2026) found that gains on embodied-reasoning benchmarks, EB-ALFRED and EB-Habitat among them, did not carry over to closed-loop robot control in SimplerEnv, which is also simulation.
Our assessment Opinion
A score reflects both the model and the framework around it. Compare models only under the same framework.
Reasoning
An EmbodiedBench score measures a whole agent. The agent is the model together with the authors' prompt, ten example plans and the simulator's ready-made skills. Changing only the agent framework moved GPT-4o from 56.3 to 75.0 on EB-ALFRED. Compare models only under the same framework and code version.
Confidence: high
Use the navigation and arm tasks to judge how much a model uses vision.
Reasoning
EB-ALFRED and EB-Habitat scores say more about planning from text than about seeing. GPT-4o did as well on them without images. Use EB-Navigation and EB-Manipulation to judge how much a model uses vision.
Confidence: medium
Training on public test trajectories can inflate scores.
Reasoning
Treat high scores from models trained on EmbodiedBench trajectories with caution. The test tasks are public, and the challenge's held-out stage showed large drops, although its navigation tasks were also harder.
Confidence: medium
The benchmark runs only in simulation. It gives no evidence about real robots.
Reasoning
None of these results show real-robot ability. Actions are high-level skills or discretised arm poses carried out by a motion planner, and the benchmark has never been compared with real robots.
Confidence: high
Known problems 6
The code has no licence
The code has no licence. A user asked for clarification and has had no reply.101213+1
Details
The repository has no LICENSE file and the README states no licence (GitHub API, 2026-10-10). On 2026-09-24 a group preparing an ICLR submission asked whether running the code and reporting scores is permitted; there is no reply yet (issue #47). Only the EB-ALFRED data card states a licence (Apache-2.0). EB-Manipulation also requires CoppeliaSim, which needs a paid licence for commercial use.
Papers average different subsets, and some published numbers do not reproduce
Papers average over 2, 4, 5 or 6 subsets, so their numbers are rarely comparable.257+4
Details
The original paper averages six subsets (five for navigation and manipulation). ERA reports five-subset averages for EB-ALFRED, which moves Claude-3.5-Sonnet from 64.0 to 66.4; Mimir averages four subsets; RoboMemory-style comparisons average only base and long-horizon. The authors' Qwen2.5-VL scores came from Alibaba's API, where some requests were blocked; they advise relying on self-hosted runs (issues #29 and #33; one user got 0.52 on EB-ALFRED long-horizon for Qwen2.5-VL-72B against 0.34 published). The paper's text names Claude-3.5-Sonnet best on EB-ALFRED at 64.0 while its Table 2 shows Claude-3.7-Sonnet at 67.7, and the project page still says 13 models were evaluated where the paper says 24.
The agent framework changes scores a lot
A score depends on the agent framework (the prompts, examples and add-ons around the model). A memory add-on raised GPT-4o from 56.3 to 75.0 on EB-ALFRED.672+1
Details
A score measures a model inside an agent framework. Adding a memory module raised GPT-4o on EB-ALFRED from 56.3 to 75.0 (BrainMem). Mimir raised InternVL3-8B from 17.5 to 60.0 (four-subset average). In the paper, removing environment feedback cost GPT-4o about 10 points on the EB-ALFRED base subset, and using no in-context examples dropped success to about 40%. Chat-history settings move EB-Navigation results by about 10 points in either direction.
Fixes in 2026 changed the test
Fixes in 2026 changed EB-ALFRED and EB-Habitat. Scores from before the fixes used the old environments.111920+1
Details
Three fixes changed results. On 2026-03-26, EB-Habitat long-horizon instructions that named the wrong starting places for objects were rewritten and the data replaced in place (issues #28 and #39). On 2026-05-30, EB-ALFRED's 'find' action was fixed: it could teleport the agent out of the room or report success when the target was not visible (issue #43, reported 2026-05-06; the fix was held back until the challenge's first stage ended). Split instructions with ambiguous 'bottle' wording were corrected the same day. Earlier results need commit 5fdee379 to reproduce. The official leaderboard still shows the paper's pre-fix numbers.
Test tasks and solved trajectories are public
Solved trajectories for the test tasks are public. The winner of the CVPR 2026 challenge scored 82.5 on the public tasks and 52.2 on held-out tasks.22115+2
Details
There is no hidden test set. The authors released trajectory datasets recorded on the benchmark's own tasks, with success flags, and suggest training on the base subset and testing on others. ERA, by the same team, trains on trajectories from three test subsets it calls 'seen' and includes them in its averages. EB-ALFRED tasks come from ALFRED's 'valid seen' split, whose rooms also appear in ALFRED's training data. In the CVPR 2026 challenge, the winning team scored 82.50 on the public tasks and 52.2 on organiser-run held-out tasks (EB-Navigation 81.67 to 37.0), though the organisers say the held-out navigation tasks were harder, and they raised the step limit to 30.
The household tasks barely need the images
In the paper's own test, GPT-4o scored 58.0 without images and 56.3 with images on EB-ALFRED, and 56.0 without and 59.0 with on EB-Habitat. Removing images hurt the navigation and manipulation tasks much more.219
Details
In the paper's own test, GPT-4o without images scored 58.0 on EB-ALFRED against 56.3 with images, and 56.0 against 59.0 on EB-Habitat; GPT-4o-mini did better without images on both (31.3 vs 24.0 and 36.7 vs 32.7). Removing images hurt the low-level environments: GPT-4o fell from 57.7 to 17.4 on EB-Navigation and from 28.9 to 16.2 on EB-Manipulation. The authors conclude the high-level tasks may rely more on text than on vision. In EB-ALFRED, the 'find' skill moves the agent to the named object without the model needing to see it.
Details
About
- What it is
- A benchmark of simulated tasks for AI agents12
More
The paper presents a fixed set of 1,128 test tasks in four environments, an agent framework and a success metric.
- Built by
- University of Illinois Urbana-Champaign, Northwestern University, University of Toronto, Toyota Technological Institute at Chicago217
More
Senior authors include Heng Ji, Huan Zhang and Tong Zhang (UIUC) and Manling Li (Northwestern).
University of Illinois Urbana-Champaign · Lead institution: 8 of 13 authors, including all three corresponding authors; two more authors interned there.2
Northwestern University · Kangrui Wang, Qineng Wang, Manling Li.2
University of Toronto · Mark Zhao (work done during an internship at UIUC).2
Toyota Technological Institute at Chicago · Marziyeh Movahedi (work done during an internship at UIUC).2
- Released
- February 2025, at ICML 202511017
More
Repository created 2025-02-12; arXiv v1 2025-02-13. Accepted at ICML 2025 (oral, per the project site).
- Version
- Paper version 3, with four environments1112
More
Paper v3 (2025-06-05) covers 24 models. The code has no tags; results from before the 2026-05-30 EB-ALFRED fixes need commit 5fdee379 to reproduce.
EB-ALFRED · 300 tasks (6 subsets x 50) from ALFRED in AI2-THOR, using Lota-Bench's simulator code. 8 high-level skill types; 171 to 298 possible actions per task.2
EB-Habitat · 300 tasks (6 x 50) from the Language Rearrangement benchmark in Habitat 2.0. 70 high-level skills; 282 instruction templates.2
EB-Navigation · 300 tasks (5 subsets x 60) in AI2-THOR, built from 90 tasks (one per scene). Low-level moves, turns and camera tilts.2
EB-Manipulation · 228 tasks (48 per subset; visual appearance 36) extending VLMBench in CoppeliaSim. 7-number arm commands, discretised.2
Setup
- Runs in
- Simulation2
More
The paper's Limitations section: evaluation is 'solely in simulated environments, without real-world experiments'.
- Simulator
- AI2-THOR, Habitat 2.0 and CoppeliaSim211
More
AI2-THOR (EB-ALFRED via Lota-Bench's code, and EB-Navigation); Habitat 2.0 (EB-Habitat); CoppeliaSim 4.1.0 through VLMBench and PyRep (EB-Manipulation)
- Robot
- Game character, Arm on wheels, Wheels only, One arm Inferred23
More
EB-ALFRED: AI2-THOR agent with abstract skills (virtual agent, as in the ALFRED entry). EB-Habitat: Fetch robot. EB-Navigation: camera agent that only moves (mobile base). EB-Manipulation: one Franka arm. The mapping is ours.
- Robot model
- Fetch and Franka Panda (simulated)32
More
EB-Habitat: Fetch robot with a suction gripper (config 'FetchSuctionRobot'). EB-Manipulation: 7-DoF Franka Emika Panda arm. EB-ALFRED and EB-Navigation: the AI2-THOR agent, no named robot. All simulated.
- Setting
- Whole home, Kitchen, Tabletop2
More
AI2-THOR rooms ('kitchens, living rooms, and bedrooms'), a ReplicaCAD apartment in Habitat, and a table with objects for EB-Manipulation.
- Tasks
- 1,128 tasks in 4 environments12
More
1,128 test tasks: EB-ALFRED 300, EB-Habitat 300, EB-Navigation 300, EB-Manipulation 228
EB-ALFRED tasks come from ALFRED's 'valid seen' split, following Lota-Bench.
- Scenes
- 90 AI2-THOR scenes in EB-Navigation2
More
EB-Navigation draws on 90 AI2-THOR scenes (one base task per scene). Scene counts for the other environments are not stated.
- Training data
- There is no training split. Trajectories recorded on the test tasks have been released.1122
More
No training split. On 2025-06-03 the authors released trajectory datasets recorded by several models on the benchmark's own tasks, with success flags, and suggest training on the base subset and testing on the others.
- Changes at test
- Instruction wording, rooms and objects vary. Inferred2
More
EB-Habitat's base subset merges Language Rearrangement's 'new scenes', 'novel objects' and 'instruction rephrasing' sets; the common-sense and complex-instruction subsets vary wording. Models are tested without task-specific training, using 10 in-context examples; EmbodiedBench defines no training split.
Scoring and access
- Scored by
- Success rate2
More
Task success judged by the simulator (PDDL goal checks in EB-ALFRED and EB-Habitat; reaching within a set distance of the target in EB-Navigation).
- Score
- Success rate, averaged within each environment2
More
Share of tasks completed in each environment, shown per capability subset and as an environment average. High-level environments give the model a menu of ready-made skills (for example 'find an apple'); EB-Manipulation's arm commands are executed by a motion planner, with detection boxes and object positions supplied. There is no official overall score across environments.
The abstract's '28.9% on average' for GPT-4o is its EB-Manipulation average, not a four-environment average.
- Trials
- One run per task2
More
Each task is run once at temperature 0, with 10 in-context examples, 500x500 images and step limits of 30 (high-level), 20 (navigation) and 15 (manipulation). Episodes also stop after more than 10 invalid actions or an empty plan.
Chat history in EB-Navigation · Results use chat history by default; turning it off moves scores by about 10 points up or down depending on the model (paper Table 10; maintainer in issue #7).218
Habitat start states · Initial states vary between runs even with a fixed seed; the maintainers say some randomness is intended (issue #37, open).24
- Who runs it
- Both teams and organisers Inferred1795
More
The official leaderboard holds the authors' own runs of 24 models. Other papers self-report. The CVPR 2026 challenge's final stage was run by the organisers on held-out tasks.
- Error bars
- Not reported Inferred221
More
Single runs; no error bars in the paper or on the leaderboard. With 300 tasks, a 95% interval is about ±5.7 points at 50% success; a 50-task subset moves in steps of 2 points (our calculation).
- Leaderboard
- Official. It lists 24 models, all with results from the paper.17219
More
Project-site leaderboard: 24 models per environment, all from the paper (data files last modified 2026-06-02). Separate challenge boards for the open split and the held-out stage.
The leaderboard values equal paper Tables 2 and 3, so they were computed before the 2026 environment fixes (our comparison).
- Code licence
- No licence is stated.101112
More
No LICENSE file and no licence statement in the README. A user asked for clarification on 2026-09-24 (issue #47, no reply yet).
The GitHub licence endpoint returns 404. Without a licence, default copyright applies (our reading; not legal advice).
- Data licence
- Apache-2.0 for EB-ALFRED. The other datasets state no licence.132522
More
EB-ALFRED dataset card: Apache-2.0 (derived from ALFRED). EB-Manipulation card and the four trajectory datasets: no licence.
- Asset licence
- Mixed. CoppeliaSim needs a paid licence for commercial use. Inferred262728+3
More
The environments build on third-party simulators and data with their own terms.
ReplicaCAD and YCB terms were not checked. Not legal advice.
AI2-THOR: Apache-2.026
ALFRED: MIT27
habitat-lab: MIT28
Lota-Bench code (LLMTaskPlanning): no licence · The repository whose simulator code EB-ALFRED builds on has no licence file.29
CoppeliaSim: commercial licence needed · The README installs CoppeliaSim Pro 4.1.0. Coppelia Robotics lists commercial use under paid licences; its free Edu edition is limited to schools and universities.1411
- Access
- Open. The code is on GitHub and the data is on Hugging Face.1113
More
Code on GitHub; datasets on Hugging Face without a gate. EB-Manipulation needs CoppeliaSim; EB-Habitat needs Habitat's YCB and ReplicaCAD downloads.
Sources 30
- 1EmbodiedBench (arXiv abstract page, v3)Paper · Feb 2025 · checked 10 Oct 2026
- 2EmbodiedBench full text v3 (Sections 3 to 6, Tables 2, 3 and 10, Appendix C, Limitations)Paper · Jun 2025 · checked 10 Oct 2026
- 3EB-Habitat task config (articulated_agent_type: FetchSuctionRobot)Repository · 2025 · checked 10 Oct 2026
- 4Vlaser (Table 1; Section 3.2)Paper · Oct 2025 · checked 10 Oct 2026
- 5ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning (Table 3)Paper · Oct 2025 · checked 10 Oct 2026
- 6BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning (Tables 1 and 3)Paper · Mar 2026 · checked 10 Oct 2026
- 7Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents (Tables 1 and 2)Paper · Aug 2026 · checked 10 Oct 2026
- 8Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI (Table 1)Paper · Sep 2025 · checked 10 Oct 2026
- 9EmbodiedBench Challenge @ CVPR 2026 (rules, open and held-out leaderboards)Official site · May 2026 · checked 10 Oct 2026
- 10GitHub API: EmbodiedBench/EmbodiedBench (stars, forks, created, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
- 11EmbodiedBench README (news, update notes, installation)Repository · 30 May 2026 · checked 10 Oct 2026
- 12EmbodiedBench issue #47: License clarificationRepository · 24 Sep 2026 · checked 10 Oct 2026
- 13Hugging Face API: EmbodiedBench datasets (licences, downloads)Index · 10 Oct 2026 · checked 10 Oct 2026
- 14Coppelia Robotics site (editions and commercial licensing)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 15EmbodiedBench issue #29: Unable to reproduce scores for Qwen2.5-VL-7B-InstructRepository · 27 Oct 2025 · checked 10 Oct 2026
- 16EmbodiedBench issue #33: higher task success for Qwen2.5-VL-72B on EB-ALFRED long-horizonRepository · 20 Nov 2025 · checked 10 Oct 2026
- 17EmbodiedBench project siteOfficial site · 2026 · checked 10 Oct 2026
- 18EmbodiedBench issue #7: Test result (chat history setting)Repository · 15 Mar 2025 · checked 10 Oct 2026
- 19EmbodiedBench issue #43: find can report success when the target is not visibleRepository · 6 May 2026 · checked 10 Oct 2026
- 20EmbodiedBench issue #39: instruction and scene mismatch in EB-Habitat long-horizon tasksRepository · 25 Mar 2026 · checked 10 Oct 2026
- 21Official leaderboard data: EB-ALFREDLeaderboard · 2 Jun 2026 · checked 10 Oct 2026
- 22EB-Alfred trajectory dataset cardRepository · 4 Jun 2025 · checked 10 Oct 2026
- 23EmbodiedBench commit historyRepository · 30 May 2026 · checked 10 Oct 2026
- 24EmbodiedBench issue #37: non-deterministic initial states in HabitatRepository · 10 Feb 2026 · checked 10 Oct 2026
- 25EB-ALFRED dataset card (Apache-2.0)Repository · 22 Feb 2025 · checked 10 Oct 2026
- 26AI2-THOR repository (Apache-2.0)Repository · 2026 · checked 10 Oct 2026
- 27ALFRED repository (MIT)Repository · 2020 · checked 10 Oct 2026
- 28habitat-lab repository (MIT)Repository · 2026 · checked 10 Oct 2026
- 29LLMTaskPlanning (Lota-Bench) repository (no licence)Repository · 2024 · checked 10 Oct 2026
- 30Semantic Scholar API record for arXiv:2502.09560Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: Paper (Limitations and full text), README, project site, challenge page; Vlaser (2510.11027); web searches on 2026-10-10 for real-robot evaluations of EmbodiedBench tasks or correlations with real-robot results. None found.
human_baseline: Paper full text, README, project site. None.
license_code: GitHub licence API (404), repository root listing (no LICENSE), README, issue #47.
industry_use: Full texts of HY-Embodied-0.5, Seed1.8, Qwen3-VL, InternVL3.5, GLM-4.5V, Xiaomi-Robotics-0, Wall-OSS-0.5 (no EmbodiedBench); web search for Pelican-VL, RoboBrain 2.5, MiMo-Embodied, Embodied-R1.5, EmbodiedBrain (no EmbodiedBench results in retrieved text); Athena-Brain report (no mention).
top_score: Official leaderboard JSON files; ERA, BrainMem, Mimir, Vlaser; challenge page. Higher numbers exist only under other protocols (items).
license_assets (ReplicaCAD, YCB): Not checked; VLMBench repository not found at the path tried.
Change history
- Basic entry created (phase 1 re-verification).
- Full entry written from primary sources, starting from the basic entry and the core-sim-b inventory record; every fact re-checked. Added environment details, robot models (Fetch from the EB-Habitat config; Franka Panda), the vision ablation, 2026 fixes and GitHub issues, the CVPR 2026 challenge open vs held-out results, later protocol variants (ERA, BrainMem, Mimir), trajectory datasets and licence findings. Confirmed prior leads: 1,128 tasks, 24 models, ICML 2025, no LICENSE file, Apache-2.0 for EB-ALFRED only, 256 citations, 347 stars, the 13-vs-24 model count conflict on the project page. Basic entry's evaluator value (null) replaced by 'both'.
- Published as a full entry.