EmbodiedBench

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

How to read this picture

EmbodiedBench is a set of 1,128 simulated tasks in which a multimodal AI model (one that takes in both images and text) plans a robot's actions. It reports the share of tasks the model completes in each of its four environments.12

Sources
Last checked 10 Oct 2026Full entry48 of 60 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Household tasks
Robot
Game character, Arm on wheels, Wheels only, One arm23
Used by
Academic agent papers. We found no company reports that use it.456+3
256 citations
Licence
No licence is stated.101112
Commercial use: unclear

What a score here does not tell you Inferred

  1. How well a model will do on a real robot.EmbodiedBench runs only in simulation. No study has compared its scores with real-robot results.
  2. Whether a model can control a robot's fine movements.The model picks from ready-made skills or gives arm poses that a motion planner carries out.
  3. How good the model is on its own.Scores change with the prompts, the examples and any memory add-ons used around the model.

Comparisons with real robots

Not checked Inferred2119+1

Details

The paper's Limitations section says evaluation is solely in simulation. Vlaser (ICLR 2026) found that gains on embodied-reasoning benchmarks, EB-ALFRED and EB-Habitat among them, did not carry over to closed-loop robot control in SimplerEnv, which is also simulation.

Our assessment Opinion

A score reflects both the model and the framework around it. Compare models only under the same framework.

Reasoning

An EmbodiedBench score measures a whole agent. The agent is the model together with the authors' prompt, ten example plans and the simulator's ready-made skills. Changing only the agent framework moved GPT-4o from 56.3 to 75.0 on EB-ALFRED. Compare models only under the same framework and code version.

Confidence: high

Use the navigation and arm tasks to judge how much a model uses vision.

Reasoning

EB-ALFRED and EB-Habitat scores say more about planning from text than about seeing. GPT-4o did as well on them without images. Use EB-Navigation and EB-Manipulation to judge how much a model uses vision.

Confidence: medium

Training on public test trajectories can inflate scores.

Reasoning

Treat high scores from models trained on EmbodiedBench trajectories with caution. The test tasks are public, and the challenge's held-out stage showed large drops, although its navigation tasks were also harder.

Confidence: medium

The benchmark runs only in simulation. It gives no evidence about real robots.

Reasoning

None of these results show real-robot ability. Actions are high-level skills or discretised arm poses carried out by a motion planner, and the benchmark has never been compared with real robots.

Confidence: high

Known problems 6

  1. The code has no licence

    The code has no licence. A user asked for clarification and has had no reply.101213+1

    Details

    The repository has no LICENSE file and the README states no licence (GitHub API, 2026-10-10). On 2026-09-24 a group preparing an ICLR submission asked whether running the code and reporting scores is permitted; there is no reply yet (issue #47). Only the EB-ALFRED data card states a licence (Apache-2.0). EB-Manipulation also requires CoppeliaSim, which needs a paid licence for commercial use.

  2. Papers average different subsets, and some published numbers do not reproduce

    Papers average over 2, 4, 5 or 6 subsets, so their numbers are rarely comparable.257+4

    Details

    The original paper averages six subsets (five for navigation and manipulation). ERA reports five-subset averages for EB-ALFRED, which moves Claude-3.5-Sonnet from 64.0 to 66.4; Mimir averages four subsets; RoboMemory-style comparisons average only base and long-horizon. The authors' Qwen2.5-VL scores came from Alibaba's API, where some requests were blocked; they advise relying on self-hosted runs (issues #29 and #33; one user got 0.52 on EB-ALFRED long-horizon for Qwen2.5-VL-72B against 0.34 published). The paper's text names Claude-3.5-Sonnet best on EB-ALFRED at 64.0 while its Table 2 shows Claude-3.7-Sonnet at 67.7, and the project page still says 13 models were evaluated where the paper says 24.

  3. The agent framework changes scores a lot

    A score depends on the agent framework (the prompts, examples and add-ons around the model). A memory add-on raised GPT-4o from 56.3 to 75.0 on EB-ALFRED.672+1

    Details

    A score measures a model inside an agent framework. Adding a memory module raised GPT-4o on EB-ALFRED from 56.3 to 75.0 (BrainMem). Mimir raised InternVL3-8B from 17.5 to 60.0 (four-subset average). In the paper, removing environment feedback cost GPT-4o about 10 points on the EB-ALFRED base subset, and using no in-context examples dropped success to about 40%. Chat-history settings move EB-Navigation results by about 10 points in either direction.

  4. Fixes in 2026 changed the test

    Fixes in 2026 changed EB-ALFRED and EB-Habitat. Scores from before the fixes used the old environments.111920+1

    Details

    Three fixes changed results. On 2026-03-26, EB-Habitat long-horizon instructions that named the wrong starting places for objects were rewritten and the data replaced in place (issues #28 and #39). On 2026-05-30, EB-ALFRED's 'find' action was fixed: it could teleport the agent out of the room or report success when the target was not visible (issue #43, reported 2026-05-06; the fix was held back until the challenge's first stage ended). Split instructions with ambiguous 'bottle' wording were corrected the same day. Earlier results need commit 5fdee379 to reproduce. The official leaderboard still shows the paper's pre-fix numbers.

  5. Test tasks and solved trajectories are public

    Solved trajectories for the test tasks are public. The winner of the CVPR 2026 challenge scored 82.5 on the public tasks and 52.2 on held-out tasks.22115+2

    Details

    There is no hidden test set. The authors released trajectory datasets recorded on the benchmark's own tasks, with success flags, and suggest training on the base subset and testing on others. ERA, by the same team, trains on trajectories from three test subsets it calls 'seen' and includes them in its averages. EB-ALFRED tasks come from ALFRED's 'valid seen' split, whose rooms also appear in ALFRED's training data. In the CVPR 2026 challenge, the winning team scored 82.50 on the public tasks and 52.2 on organiser-run held-out tasks (EB-Navigation 81.67 to 37.0), though the organisers say the held-out navigation tasks were harder, and they raised the step limit to 30.

  6. The household tasks barely need the images

    In the paper's own test, GPT-4o scored 58.0 without images and 56.3 with images on EB-ALFRED, and 56.0 without and 59.0 with on EB-Habitat. Removing images hurt the navigation and manipulation tasks much more.219

    Details

    In the paper's own test, GPT-4o without images scored 58.0 on EB-ALFRED against 56.3 with images, and 56.0 against 59.0 on EB-Habitat; GPT-4o-mini did better without images on both (31.3 vs 24.0 and 36.7 vs 32.7). Removing images hurt the low-level environments: GPT-4o fell from 57.7 to 17.4 on EB-Navigation and from 28.9 to 16.2 on EB-Manipulation. The authors conclude the high-level tasks may rely more on text than on vision. In EB-ALFRED, the 'find' skill moves the agent to the named object without the model needing to see it.

Details

About

What it is
A benchmark of simulated tasks for AI agents12
More

The paper presents a fixed set of 1,128 test tasks in four environments, an agent framework and a success metric.

Built by
University of Illinois Urbana-Champaign, Northwestern University, University of Toronto, Toyota Technological Institute at Chicago217
More

Senior authors include Heng Ji, Huan Zhang and Tong Zhang (UIUC) and Manling Li (Northwestern).

University of Illinois Urbana-Champaign · Lead institution: 8 of 13 authors, including all three corresponding authors; two more authors interned there.2

Northwestern University · Kangrui Wang, Qineng Wang, Manling Li.2

University of Toronto · Mark Zhao (work done during an internship at UIUC).2

Toyota Technological Institute at Chicago · Marziyeh Movahedi (work done during an internship at UIUC).2

Released
February 2025, at ICML 202511017
More

Repository created 2025-02-12; arXiv v1 2025-02-13. Accepted at ICML 2025 (oral, per the project site).

Version
Paper version 3, with four environments1112
More

Paper v3 (2025-06-05) covers 24 models. The code has no tags; results from before the 2026-05-30 EB-ALFRED fixes need commit 5fdee379 to reproduce.

EB-ALFRED · 300 tasks (6 subsets x 50) from ALFRED in AI2-THOR, using Lota-Bench's simulator code. 8 high-level skill types; 171 to 298 possible actions per task.2

EB-Habitat · 300 tasks (6 x 50) from the Language Rearrangement benchmark in Habitat 2.0. 70 high-level skills; 282 instruction templates.2

EB-Navigation · 300 tasks (5 subsets x 60) in AI2-THOR, built from 90 tasks (one per scene). Low-level moves, turns and camera tilts.2

EB-Manipulation · 228 tasks (48 per subset; visual appearance 36) extending VLMBench in CoppeliaSim. 7-number arm commands, discretised.2

Last update
May 2026, with fixes to the environments2311
More

2026-05-30: two EB-ALFRED fixes (the find action and ambiguous split instructions). Earlier in 2026: EB-Habitat long-horizon data replaced in place (2026-03-26) and Qwen3-VL support (2026-04-08).

Status
Active. The last change was in May 2026. Inferred23912
More

Commits through 2026-05-30; the CVPR 2026 challenge ran from April to May 2026. A licence question opened 2026-09-24 (issue #47) has no reply yet.

Last commit about four months before 2026-10-10.

Setup

Runs in
Simulation2
More

The paper's Limitations section: evaluation is 'solely in simulated environments, without real-world experiments'.

Simulator
AI2-THOR, Habitat 2.0 and CoppeliaSim211
More

AI2-THOR (EB-ALFRED via Lota-Bench's code, and EB-Navigation); Habitat 2.0 (EB-Habitat); CoppeliaSim 4.1.0 through VLMBench and PyRep (EB-Manipulation)

Robot
Game character, Arm on wheels, Wheels only, One arm Inferred23
More

EB-ALFRED: AI2-THOR agent with abstract skills (virtual agent, as in the ALFRED entry). EB-Habitat: Fetch robot. EB-Navigation: camera agent that only moves (mobile base). EB-Manipulation: one Franka arm. The mapping is ours.

Robot model
Fetch and Franka Panda (simulated)32
More

EB-Habitat: Fetch robot with a suction gripper (config 'FetchSuctionRobot'). EB-Manipulation: 7-DoF Franka Emika Panda arm. EB-ALFRED and EB-Navigation: the AI2-THOR agent, no named robot. All simulated.

Setting
Whole home, Kitchen, Tabletop2
More

AI2-THOR rooms ('kitchens, living rooms, and bedrooms'), a ReplicaCAD apartment in Habitat, and a table with objects for EB-Manipulation.

Tasks
1,128 tasks in 4 environments12
More

1,128 test tasks: EB-ALFRED 300, EB-Habitat 300, EB-Navigation 300, EB-Manipulation 228

EB-ALFRED tasks come from ALFRED's 'valid seen' split, following Lota-Bench.

Scenes
90 AI2-THOR scenes in EB-Navigation2
More

EB-Navigation draws on 90 AI2-THOR scenes (one base task per scene). Scene counts for the other environments are not stated.

Training data
There is no training split. Trajectories recorded on the test tasks have been released.1122
More

No training split. On 2025-06-03 the authors released trajectory datasets recorded by several models on the benchmark's own tasks, with success flags, and suggest training on the base subset and testing on the others.

Changes at test
Instruction wording, rooms and objects vary. Inferred2
More

EB-Habitat's base subset merges Language Rearrangement's 'new scenes', 'novel objects' and 'instruction rephrasing' sets; the common-sense and complex-instruction subsets vary wording. Models are tested without task-specific training, using 10 in-context examples; EmbodiedBench defines no training split.

Scoring and access

Scored by
Success rate2
More

Task success judged by the simulator (PDDL goal checks in EB-ALFRED and EB-Habitat; reaching within a set distance of the target in EB-Navigation).

Score
Success rate, averaged within each environment2
More

Share of tasks completed in each environment, shown per capability subset and as an environment average. High-level environments give the model a menu of ready-made skills (for example 'find an apple'); EB-Manipulation's arm commands are executed by a motion planner, with detection boxes and object positions supplied. There is no official overall score across environments.

The abstract's '28.9% on average' for GPT-4o is its EB-Manipulation average, not a four-environment average.

Trials
One run per task2
More

Each task is run once at temperature 0, with 10 in-context examples, 500x500 images and step limits of 30 (high-level), 20 (navigation) and 15 (manipulation). Episodes also stop after more than 10 invalid actions or an empty plan.

Chat history in EB-Navigation · Results use chat history by default; turning it off moves scores by about 10 points up or down depending on the model (paper Table 10; maintainer in issue #7).218

Habitat start states · Initial states vary between runs even with a fixed seed; the maintainers say some randomness is intended (issue #37, open).24

Who runs it
Both teams and organisers Inferred1795
More

The official leaderboard holds the authors' own runs of 24 models. Other papers self-report. The CVPR 2026 challenge's final stage was run by the organisers on held-out tasks.

Error bars
Not reported Inferred221
More

Single runs; no error bars in the paper or on the leaderboard. With 300 tasks, a 95% interval is about ±5.7 points at 50% success; a 50-task subset moves in steps of 2 points (our calculation).

Leaderboard
Official. It lists 24 models, all with results from the paper.17219
More

Project-site leaderboard: 24 models per environment, all from the paper (data files last modified 2026-06-02). Separate challenge boards for the open split and the held-out stage.

The leaderboard values equal paper Tables 2 and 3, so they were computed before the 2026 environment fixes (our comparison).

Code licence
No licence is stated.101112
More

No LICENSE file and no licence statement in the README. A user asked for clarification on 2026-09-24 (issue #47, no reply yet).

The GitHub licence endpoint returns 404. Without a licence, default copyright applies (our reading; not legal advice).

Data licence
Apache-2.0 for EB-ALFRED. The other datasets state no licence.132522
More

EB-ALFRED dataset card: Apache-2.0 (derived from ALFRED). EB-Manipulation card and the four trajectory datasets: no licence.

Asset licence
Mixed. CoppeliaSim needs a paid licence for commercial use. Inferred262728+3
More

The environments build on third-party simulators and data with their own terms.

ReplicaCAD and YCB terms were not checked. Not legal advice.

AI2-THOR: Apache-2.026

ALFRED: MIT27

habitat-lab: MIT28

Lota-Bench code (LLMTaskPlanning): no licence · The repository whose simulator code EB-ALFRED builds on has no licence file.29

CoppeliaSim: commercial licence needed · The README installs CoppeliaSim Pro 4.1.0. Coppelia Robotics lists commercial use under paid licences; its free Edu edition is limited to schools and universities.1411

Access
Open. The code is on GitHub and the data is on Hugging Face.1113
More

Code on GitHub; datasets on Hugging Face without a gate. EB-Manipulation needs CoppeliaSim; EB-Habitat needs Habitat's YCB and ReplicaCAD downloads.

Commercial use
Unclear Inferred102514
More

The code has no licence, so no rights are granted by default (our reading). EB-ALFRED data is Apache-2.0; other data has no licence. EB-Manipulation needs CoppeliaSim, which requires a paid licence for commercial use. Not legal advice.

Published at
ICML 2025, as an oral presentation11730
More

ICML 2025 (arXiv comment 'Accepted to ICML 2025'; project site says oral)

Sources 30

  1. 1EmbodiedBench (arXiv abstract page, v3)Paper · Feb 2025 · checked 10 Oct 2026
  2. 2EmbodiedBench full text v3 (Sections 3 to 6, Tables 2, 3 and 10, Appendix C, Limitations)Paper · Jun 2025 · checked 10 Oct 2026
  3. 3EB-Habitat task config (articulated_agent_type: FetchSuctionRobot)Repository · 2025 · checked 10 Oct 2026
  4. 4Vlaser (Table 1; Section 3.2)Paper · Oct 2025 · checked 10 Oct 2026
  5. 5ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning (Table 3)Paper · Oct 2025 · checked 10 Oct 2026
  6. 6BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning (Tables 1 and 3)Paper · Mar 2026 · checked 10 Oct 2026
  7. 7Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents (Tables 1 and 2)Paper · Aug 2026 · checked 10 Oct 2026
  8. 8Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI (Table 1)Paper · Sep 2025 · checked 10 Oct 2026
  9. 9EmbodiedBench Challenge @ CVPR 2026 (rules, open and held-out leaderboards)Official site · May 2026 · checked 10 Oct 2026
  10. 10GitHub API: EmbodiedBench/EmbodiedBench (stars, forks, created, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
  11. 11EmbodiedBench README (news, update notes, installation)Repository · 30 May 2026 · checked 10 Oct 2026
  12. 12EmbodiedBench issue #47: License clarificationRepository · 24 Sep 2026 · checked 10 Oct 2026
  13. 13Hugging Face API: EmbodiedBench datasets (licences, downloads)Index · 10 Oct 2026 · checked 10 Oct 2026
  14. 14Coppelia Robotics site (editions and commercial licensing)Official site · 10 Oct 2026 · checked 10 Oct 2026
  15. 15EmbodiedBench issue #29: Unable to reproduce scores for Qwen2.5-VL-7B-InstructRepository · 27 Oct 2025 · checked 10 Oct 2026
  16. 16EmbodiedBench issue #33: higher task success for Qwen2.5-VL-72B on EB-ALFRED long-horizonRepository · 20 Nov 2025 · checked 10 Oct 2026
  17. 17EmbodiedBench project siteOfficial site · 2026 · checked 10 Oct 2026
  18. 18EmbodiedBench issue #7: Test result (chat history setting)Repository · 15 Mar 2025 · checked 10 Oct 2026
  19. 19EmbodiedBench issue #43: find can report success when the target is not visibleRepository · 6 May 2026 · checked 10 Oct 2026
  20. 20EmbodiedBench issue #39: instruction and scene mismatch in EB-Habitat long-horizon tasksRepository · 25 Mar 2026 · checked 10 Oct 2026
  21. 21Official leaderboard data: EB-ALFREDLeaderboard · 2 Jun 2026 · checked 10 Oct 2026
  22. 22EB-Alfred trajectory dataset cardRepository · 4 Jun 2025 · checked 10 Oct 2026
  23. 23EmbodiedBench commit historyRepository · 30 May 2026 · checked 10 Oct 2026
  24. 24EmbodiedBench issue #37: non-deterministic initial states in HabitatRepository · 10 Feb 2026 · checked 10 Oct 2026
  25. 25EB-ALFRED dataset card (Apache-2.0)Repository · 22 Feb 2025 · checked 10 Oct 2026
  26. 26AI2-THOR repository (Apache-2.0)Repository · 2026 · checked 10 Oct 2026
  27. 27ALFRED repository (MIT)Repository · 2020 · checked 10 Oct 2026
  28. 28habitat-lab repository (MIT)Repository · 2026 · checked 10 Oct 2026
  29. 29LLMTaskPlanning (Lota-Bench) repository (no licence)Repository · 2024 · checked 10 Oct 2026
  30. 30Semantic Scholar API record for arXiv:2502.09560Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: Paper (Limitations and full text), README, project site, challenge page; Vlaser (2510.11027); web searches on 2026-10-10 for real-robot evaluations of EmbodiedBench tasks or correlations with real-robot results. None found.

human_baseline: Paper full text, README, project site. None.

license_code: GitHub licence API (404), repository root listing (no LICENSE), README, issue #47.

industry_use: Full texts of HY-Embodied-0.5, Seed1.8, Qwen3-VL, InternVL3.5, GLM-4.5V, Xiaomi-Robotics-0, Wall-OSS-0.5 (no EmbodiedBench); web search for Pelican-VL, RoboBrain 2.5, MiMo-Embodied, Embodied-R1.5, EmbodiedBrain (no EmbodiedBench results in retrieved text); Athena-Brain report (no mention).

top_score: Official leaderboard JSON files; ERA, BrainMem, Mimir, Vlaser; challenge page. Higher numbers exist only under other protocols (items).

license_assets (ReplicaCAD, YCB): Not checked; VLMBench repository not found at the path tried.

Change history

  1. Basic entry created (phase 1 re-verification).
  2. Full entry written from primary sources, starting from the basic entry and the core-sim-b inventory record; every fact re-checked. Added environment details, robot models (Fetch from the EB-Habitat config; Franka Panda), the vision ablation, 2026 fixes and GitHub issues, the CVPR 2026 challenge open vs held-out results, later protocol variants (ERA, BrainMem, Mimir), trajectory datasets and licence findings. Confirmed prior leads: 1,128 tasks, 24 models, ICML 2025, no LICENSE file, Apache-2.0 for EB-ALFRED only, 256 citations, 347 stars, the 13-vs-24 model count conflict on the project page. Basic entry's evaluator value (null) replaced by 'both'.
  3. Published as a full entry.