ALFRED

ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

How to read this picture

ALFRED is a benchmark in which an AI agent follows written instructions to do household tasks in simulated rooms. Agents are scored on test rooms they did not see in training.123

Sources
Last checked 10 Oct 2026Full entry67 of 79 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Household tasks
Robot
Game character23
None
Used by
95 leaderboard entries34
1,230 citations
Licence
MIT5
Commercial use: allowed

What a score here does not tell you Inferred

  1. How well an agent would do on a real robot.We found no study that compares ALFRED scores with real-robot results.
  2. Whether an agent can physically grasp objects.The agent selects an object by drawing a mask, which marks the object in the camera image.
  3. Whether an agent follows each step of the instructions.Older models ignored the step-by-step instructions.
  4. How efficiently an agent completes a task.The leaderboard ranks entries by success rate. It does not use path length.
ChartPublished scores over time
95% AND ABOVE01020304050607080901002020202120222023202420252026Seq2Seq+PM (paper baseline) · 0.39% · 2020-03ECCV 2020 winner · 4.45% · 2020-08HiTUT · 13.87% · 2021-01HLSM · 16.29% · 2021-06FILM · 26.49% · 2021-09EPA · 36.07% · 2022-05Prompter · 45.72% · 2022-08ECLAIR · 50.36% · 2023-06RoboGPT · 62% · 2023-12EPO · 62.35% · 2024-02GRL · 68.52% · 2025-07PACE-Agent · 65.53% · 2026-09Seq2Seq+PM 0.39%GRL 68.52%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.326+7

Comparisons with real robots

Not checked Inferred21415

Details

ALFRED actions are discrete steps and mask-based interactions with no robot model, so a direct robot transfer is not defined. ReALFRED (ECCV 2024) moves ALFRED-style tasks into 3D-captured real homes, still in simulation, and finds that methods built for ALFRED score lower on all metrics. RoboGPT, a top-10 leaderboard entry, reports no real-robot experiment in its v3 text. See searched.

Our assessment Opinion

A good ALFRED score is weak evidence about how an agent would do on real robots.

Reasoning

An ALFRED score shows how well an agent completes scripted household tasks from written instructions in 120 game-engine rooms. It is weak evidence about robots. Actions are discrete steps, objects are selected by drawing a mask, and no study has compared ALFRED scores with real-robot results.

Confidence: high

Do not read much into gaps of a few points between top entries.

Reasoning

Small gaps between top entries are hard to interpret. The unseen test split covers 8 rooms, and entries differ in which instructions they use. In one study, a model that led on the validation rooms scored lower on the test rooms.

Confidence: medium

Check the path-weighted scores as well. Top agents often take long routes to finish a task.

Reasoning

Read the path-weighted score (a score that gives less credit when the agent takes more steps than the expert) next to the success rate. Several top entries reach their success rate through long searches. GRL scores 68.52% success but 45.47% path-weighted, and EPA scores 36.07% against 2.92%.

Confidence: high

ALFRED is still in active use. The best entry is well below human performance.

Reasoning

ALFRED is still an active comparison point for language-driven household agents. The leaderboard received new entries in September 2026, and the best entry (68.52%) remains well below human success (91.0%).

Confidence: medium

Known problems 5

  1. Leaderboard rows use different inputs

    The leaderboard ranks entries that use only the goal together with entries that use the full instructions.3114

    Details

    The leaderboard ranks all entries by one success rate, whether or not they use the step-by-step instructions. Some rows are labelled goal-only ('HIA-High-Goal-Only', 'high level only'). Prompter's paper reports 45.32% with step-by-step instructions and 41.53% with the goal alone. Its leaderboard row of 45.72% used domain knowledge that its authors judged 'too specific to ALFRED', so the paper reports the lower 45.32% row ('Prompter, no slice replay'). Teams answer a questionnaire on their inputs, but the answers are not shown in the table.

  2. The benchmark depends on an old simulator release

    Scores depend on AI2-THOR 2.1.0, a simulator release from 2019. Inferred16174

    Details

    The code requires AI2-THOR 2.1.0 (uploaded to PyPI on 2019-09-06), PyTorch 1.1.0 and an X server for rendering. The latest AI2-THOR on PyPI is 5.0.0 (2022-12-13). Results depend on replaying actions in version 2.1.0, so new work cannot move to newer AI2-THOR releases without breaking comparability.

  3. Scores drop in more realistic scenes

    Methods built for ALFRED score lower in multi-room scenes captured in 3D.142

    Details

    ReALFRED (ECCV 2024) rebuilds ALFRED-style tasks in larger, multi-room, 3D-captured scenes. Methods designed for ALFRED consistently scored lower on all metrics there. ALFRED's own rooms are single game-engine rooms.

  4. Agents make little use of the step-by-step instructions

    Removing the step-by-step instructions barely changed the success of older models.18

    Details

    ALFRED-L (EMNLP 2022) tested six models from 2020 and 2021. Removing every step-by-step instruction and keeping only the goal lowered success by at most 6.7% in relative terms. Adding one 'go back' step to the instructions cut the success of the best model (ET+Synth) on seen rooms from 44.7% to 9.2%. The authors conclude that the models rely on the usual order in which objects are visited in ALFRED's 7 task structures. Newer modular and language-model agents were not tested.

  5. Validation gains did not carry over to the test set

    A model that led on the validation rooms fell behind on the test rooms.192

    Details

    Kim et al. (UNC Chapel Hill and Amazon Alexa AI, 2022) built a model that beat the best published models on unseen validation (13.8% success against 12.55% for ABP) but scored lower on unseen test (8.57% against 15.43%). They saw a 3-point spread in validation success across random seeds for the same model. The unseen validation split has only 4 rooms and the unseen test split 8 rooms, so model selection on one small set of rooms may not transfer to the other. The authors suggest averaging several runs of each model when ranking methods.

Details

About

What it is
Benchmark12
More

The paper and the project site present ALFRED as a benchmark with fixed splits, fixed metrics and a test server.

Built by
University of Washington, Carnegie Mellon University, Allen Institute for AI, NVIDIA2
More

Authors: Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, Dieter Fox.

University of Washington · Paul G. Allen School. Seven of the eight authors list it.2

Carnegie Mellon University · Language Technologies Institute (Yonatan Bisk).2

Allen Institute for AI · Yonatan Bisk, Winson Han, Roozbeh Mottaghi. The data bucket on Amazon S3 is named ai2-vision-alfred.220

NVIDIA · Second affiliation of Dieter Fox.2

Released
December 2019. Published at CVPR 2020.1
More

arXiv v1 on 2019-12-03. Published at CVPR 2020.

The GitHub repository was created on 2019-12-12 (GitHub API).

Version
Data 2.1.0 for AI2-THOR 2.1.042116
More

Dataset release 2.1.0 (json, json_feat and full packages) for the AI2-THOR 2.1.0 simulator. No later dataset version exists.

AI2-THOR 2.1.0 · Pinned in requirements.txt together with PyTorch 1.1.0. Uploaded to PyPI on 2019-09-06. The latest AI2-THOR on PyPI is 5.0.0 (2022-12-13).1617

Paper v2 (2020-03-31) · Results table updated after the 2020-03-28 change that selects the target object by mask overlap (IoU).14

Errata for the Goto sub-goal · The Goto sub-goal success rule changed to 'within 3 planner steps'. The main results and the leaderboard are not affected.224

Quickstart data fix (2020-10-26) · A missing stop frame was added to the json_feat package.420

Last update
September 2026. A new entry was added to the leaderboard.3234
More

Newest leaderboard entry on 2026-09-28. Last repository commit on 2026-02-05 (README only). Since April 2025, teams email their results file and the authors post the scores.

Status
The leaderboard is active. The code has not changed since 2021. Inferred32316
More

The leaderboard still receives and posts entries in 2026. The code has not changed since 2021-08 apart from README edits and one dependency pin.

Commit file lists: every commit after 2021-08-29 touches only README.md or requirements.txt (a flask version pin authored 2023-02-28, merged 2025-04-23). The Ai2-hosted leaderboard was deprecated in April 2025 and replaced by email submissions. 23 issues are open (GitHub API).

Setup

Runs in
Simulation2
Simulator
AI2-THOR 2.1.0 (Unity)216
More

AI2-THOR 2.1.0, a Unity-based household simulator from the Allen Institute for AI. The paper used AI2-THOR 2.0; the code requires 2.1.0.

Robot
Game character Inferred23
More

An egocentric agent with 5 navigation actions (move ahead, rotate left or right, look up or down), 7 interaction actions (pick up, put, open, close, toggle on, toggle off, slice) that target objects through a pixel mask, and a stop action. It has no robot body model. 'virtual-agent' is the closest taxonomy value; 'mobile-manipulator' would overstate the physics. At test time, leaderboard entries may use only RGB images and language.

Robot model
No robot model. An abstract first-person agent with discrete steps and mask-based object interaction.2
Setting
Kitchen, Whole home Inferred2
More

Room counts from the paper, Section 3. Bathrooms, bedrooms and living rooms are mapped to 'home'.

Tasks
7 task types in 2,685 task settings2
More

7 task types (for example Pick & Place, Heat & Place, Clean & Place, Examine in Light) over 2,685 combinations of object, receptacle and room

Scenes
120 rooms, including 8 unseen test rooms2
More

120 rooms in AI2-THOR. Splits by room: train 108, validation seen 88, validation unseen 4, test seen 107, test unseen 8.

Paper Table 2. Seen rooms are subsets of the training rooms.

Training data
8,055 planner demonstrations and 25,743 directives2
More

8,055 expert demonstrations (3 per task setting) generated by a PDDL planner, with 25,743 crowd-written English directives. Average 50 steps; 428,322 image-action pairs.

Directives were written on Amazon Mechanical Turk by at least three annotators per demonstration and checked by a second group of workers.

Directive splits · Train 21,023; validation seen 820; validation unseen 821; test seen 1,533; test unseen 1,529 (paper Table 2)2

Planner-made, filtered · The organisers' retrospective says task configurations the PDDL planner could not solve were dropped, which leaves fewer corner cases than in TEACh.24

Download sizes · Trajectory JSONs 35,748,602 bytes; JSONs with ResNet features 17,835,841,659 bytes; full package 116,073,781,310 bytes (HTTP Content-Length, read 2026-10-10)2021

Changes at test
New rooms, object variants and instructions Inferred2
More

The headline 'unseen' test split uses 8 rooms that do not appear in training, with new object variants and newly written instructions. The 'seen' split reuses training rooms.

Paper Table 2 and Section 3: unseen validation (4 scenes) and unseen test (8 scenes) are distinct from training and from each other; the split examines generalisation 'to entirely new spaces with novel object class variations'. The 7 task types and 84 object classes are the same in training and test.

Scoring and access

Scored by
Success rate, Progress score, Path efficiency23
More

Task success is the headline. Goal-condition success gives partial credit. Both also come in a path-length-weighted form.

Score
Success rate on unseen test rooms, with all goals met23
More

Task success: an episode counts only if every goal condition holds at the end, for example a potato slice is heated and lies on a counter. Goal-condition success: the share of goal conditions met (2.55 per task on average), which gives partial credit. Path-length-weighted (PLW) versions multiply each score by the expert's step count divided by the larger of the expert's and the agent's step counts, so taking twice as long as the expert halves the credit. The leaderboard ranks entries by success rate on the unseen test split.

Paper Section 5.1 and the leaderboard 'Scoring' section. Episodes end after 1,000 steps or 10 failed actions. Humans scored 91.0% task success and 94.5% goal-condition success on 100 unseen test directives (paper Section 6.4).

Trials
Each of the 1,529 unseen and 1,533 seen test directives is run once. Limits: 1,000 steps and 10 failed actions per episode.234
More

Teams run their agent locally on the test instructions and submit the action sequences; the server replays them in the simulator. The README forbids changing the step and failure limits.

Who runs it
The organisers run the tests34
More

The agent itself runs on the team's machine, so the organisers check input rules through a questionnaire (RGB and language only at test time) rather than by running the policy.

Error bars
Not reported Inferred319
More

The leaderboard shows single numbers. One study reports a seed-to-seed spread of 3 points in success rate for the same model on unseen validation (s20).

Leaderboard
Official and active, with 95 entries34
More

Official leaderboard on askforalfred.com: 95 entries from 2020-03-28 to 2026-09-28. One submission per 7 days; all submissions are public; email submission since April 2025.

Row count and dates from the CSV embedded in the leaderboard page, read 2026-10-10.

Code licence
MIT5
More

LICENSE file: MIT, copyright 2019 ALFRED.

Data licence
MIT (repository licence) Inferred421
More

MIT, by our reading. The README's licence section says 'MIT License'; there is no separate data licence.

The repository does not say explicitly that MIT covers the data hosted on Amazon S3. We read the repository licence as covering it. The data README and download script name no other terms.

Asset licence
The rooms and 3D objects come with the AI2-THOR simulator, whose repository and PyPI package are Apache-2.0. Inferred2517
More

We found no separate licence for the scenes and object models in AI2-THOR's builds. We read the Apache-2.0 licence of the AI2-THOR repository as covering them. Not checked: third-party asset terms inside the Unity builds.

Access
Open download2120
More

Direct downloads from Amazon S3, no registration. The three archive links answered HTTP 200 on 2026-10-10 (headers only). Test-set scores require emailing the results file.

Commercial use
Allowed Inferred5425
More

Code MIT, data MIT by our reading, simulator Apache-2.0. All three allow commercial use with attribution. Asset coverage is inferred (see license_assets). Not legal advice.

Sources 25

  1. 1ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks (arXiv abstract page)Paper · Dec 2019 · checked 10 Oct 2026
  2. 2ALFRED paper, full text v2 (Tables 2 and 3, Sections 3, 5 and 6)Paper · Mar 2020 · checked 10 Oct 2026
  3. 3ALFRED leaderboard (rules, metrics and the results table embedded as CSV)Leaderboard · Sep 2026 · checked 10 Oct 2026
  4. 4ALFRED GitHub README (leaderboard rules, April 2025 update, change log, licence)Repository · Feb 2026 · checked 10 Oct 2026
  5. 5ALFRED LICENSE fileRepository · 2019 · checked 10 Oct 2026
  6. 6ALFRED Challenge at the EVAL workshop, ECCV 2020Official site · 2020 · checked 10 Oct 2026
  7. 7ALFRED Challenge at the Embodied AI Workshop, CVPR 2021Official site · 2021 · checked 10 Oct 2026
  8. 8A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution (HLSM)Paper · Jul 2021 · checked 10 Oct 2026
  9. 9FILM: Following Instructions in Language with Modular MethodsPaper · Oct 2021 · checked 10 Oct 2026
  10. 10ALFRED Challenge at the Embodied AI Workshop, CVPR 2022Official site · 2022 · checked 10 Oct 2026
  11. 11Prompter: Utilizing Large Language Model Prompting for a Data Efficient Embodied Instruction Following (Table I and footnote 2)Paper · Nov 2022 · checked 10 Oct 2026
  12. 12ALFRED+TEACh Generalist Language Grounding Agents Challenge, CVPR 2023Official site · 2023 · checked 10 Oct 2026
  13. 13EPO: Hierarchical LLM Agents with Environment Preference OptimizationPaper · Aug 2024 · checked 10 Oct 2026
  14. 14ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments (ECCV 2024)Paper · Jul 2024 · checked 10 Oct 2026
  15. 15RoboGPT: an LLM-based Embodied Long-term Decision Making agent for Instruction Following Tasks (v3)Paper · Nov 2023 · checked 10 Oct 2026
  16. 16ALFRED requirements.txt (ai2thor==2.1.0, torch==1.1.0)Repository · Apr 2025 · checked 10 Oct 2026
  17. 17PyPI: ai2thor (release history; 2.1.0 uploaded 2019-09-06; latest 5.0.0)Repository · 13 Dec 2022 · checked 10 Oct 2026
  18. 18ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual Environments (EMNLP 2022)Paper · Dec 2022 · checked 10 Oct 2026
  19. 19On the Limits of Evaluating Embodied Agent Model Generalization Using Validation Sets (Kim et al., ACL 2022 Insights Workshop)Paper · May 2022 · checked 10 Oct 2026
  20. 20ALFRED data package on Amazon S3 (HTTP headers of json, json_feat and full 2.1.0 archives)Dataset page · 26 Oct 2020 · checked 10 Oct 2026
  21. 21ALFRED dataset README and download scriptRepository · 2020 · checked 10 Oct 2026
  22. 22ALFRED errata for Goto sub-goal evaluationRepository · 2020 · checked 10 Oct 2026
  23. 23ALFRED commit history (files changed per commit)Repository · 5 Feb 2026 · checked 10 Oct 2026
  24. 24Retrospectives on the Embodied AI Workshop (Sections 3.3.2 and 3.3.3)Paper · Oct 2022 · checked 10 Oct 2026
  25. 25AI2-THOR repository LICENSE (Apache License 2.0)Repository · 2017 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: ALFRED paper (arXiv 1912.01734v2, full text), project site and challenge pages (EVAL 2020, EAI21, EAI22, EAI23), README; web searches on 2026-10-10 for 'ALFRED real robot', 'ALFRED sim-to-real' and for real-robot experiments by top leaderboard entries; RoboGPT v3 text (no physical robot); ReALFRED (real scans, still simulated). No paired evaluation of the same agents on ALFRED and on a real robot found.

license_data: README licence section, LICENSE file, data README, download_data.sh, project site. No separate data licence found.

license_assets: AI2-THOR repository LICENSE and PyPI metadata. No separate licence for scenes or 3D objects found; third-party terms inside the Unity builds not checked.

uncertainty_reported: Leaderboard table (single numbers), paper Table 3 (single numbers), s20 (seed spread).

industry_use (frontier-model reports): Web search for ALFRED and EB-ALFRED in 2025-2026 model technical reports. None found that cites an ALFRED leaderboard score.

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
  2. Full entry. Re-checked every basic fact at the source. Corrections: capability note (manipulation is mask-based); evaluator set to organiser-run (server replay of submitted actions); sim_to_real recorded as 'none-found' (inferred) instead of unknown. Added leaderboard history from the embedded CSV (95 rows), five issues, derived benchmarks, adoption and licences for the simulator.
  3. Published as a full entry.