PARTNR

PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

How to read this picture

PARTNR is a set of 100,000 simulated household tasks in which a robot's planner (the model that decides its next steps) works with a person. It is mainly used to test task planning and coordination by large language models (LLMs).12

Sources
Last checked 10 Oct 2026Full entry44 of 56 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Working with people
Robot
Arm on wheels2
Boston Dynamics Spot (simulated)
Used by
A few third-party papers345+1
98 citations
Licence
MIT78
Commercial use: not allowed

What a score here does not tell you Inferred

  1. How well the planner will work on a real robot.We found no runs on real robots.
  2. How good the robot is at low-level manipulation (moving and handling objects).The robot's skills are either oracle skills provided by the simulator or fixed learned modules.
  3. How a planner does on the test set.The test episodes have not been released.

Comparisons with real robots

Not checked Inferred294

Details

The paper has no real-robot runs. Meta's blog calls testing in physical-world scenarios a future goal. FLEET ran two real Spot robots on different inspection tasks, with no paired PARTNR comparison.

Our assessment Opinion

PARTNR tests planners in simulation. It does not test robot control.

Reasoning

PARTNR measures high-level planning and coordination with fixed skills in simulation. A high score says little about low-level manipulation or real robots. With learned skills and realistic perception, the best LLM success fell to 0.30 (or 0.25) in the paper.

Confidence: medium

People do better than LLM agents, but part of the gap comes from differences in how they were tested.

Reasoning

The gap between people and LLM agents is real, but the headline numbers overstate it. People had retries and feedback, and the automated runs pair the LLM robot with a weaker LLM partner.

Confidence: medium

Check the split, the skills and the metric definitions before you compare numbers across papers.

Reasoning

Do not compare PARTNR numbers across papers without checking the split, the skills, the perception setting and the metric definitions. The released data supports only validation-set comparisons.

Confidence: high

Known problems 6, 1 disputed

  1. Released data does not match the paper

    The release has no test split. The weights of the fine-tuned planner are missing.21011+3

    Details

    The paper describes 100,000 train, 1,000 validation and 1,000 test episodes. The release has 111,652 people-verified train episodes (131,991 unverified) and 1,000 validation episodes, and no test split, so the paper's Table 11 cannot be reproduced. The paper says weights of the fine-tuned planner were released; the dataset holds only the skill checkpoints and an issue reporting this has had no reply since 2025-11-22. Another open issue reports that the code for task offloading and related metrics covers only the human runs. The constrained decoding library does not support Qwen or Llama-3.2 and newer models.

  2. Later papers change the metrics or the task set

    Papers by other groups redefine the metrics or use only some of the tasks.345+1

    Details

    MIT Lincoln Lab redefines success rate as the share of successful decision points and percent complete as the share of successful episodes, ran each setting for ten hours (different episode counts per model), and reprints the paper's Llama-3.1-70B rows next to its own. FLEET reports three of the four task types. AHAT scores planning only. The README notes that constrained generation, which the paper uses to block invalid actions, is not supported for OpenAI models.

  3. Some tasks are already done at the startDisputed

    Some episodes are already solved at the start. The authors say this is intended.152

    Other view: A co-author says these episodes are intentional and that there are few of them.15

    Details

    Users report validation episode 139, where the objects already met the goal and an agent succeeded by checking and calling Done. The paper's own error analysis lists 'Already Satisfied' as a task-generation failure (2% of sampled constraint-free, spatial and temporal episodes).

    A PARTNR co-author replied that such episodes are intentional, to test whether agents can verify that a task is complete, and that filtering keeps their number limited. An impossible test episode was fixed.

  4. People and LLMs are scored under different rules

    People could retry each task and got feedback after each try. LLM agents got one run.216

    Details

    The headline 'humans solve 93%, LLMs 30%' compares different set-ups. In the human runs each task could be tried up to 3 times, with a written explanation of what went wrong after each try, and the best try was kept. Tasks that no person solved in 6 tries had already been removed from the dataset. The 0.30 comes from LLM planners with learned skills, scene graphs built from camera images and an LLM-controlled partner, in one run. When the partner is a real person, the same kind of LLM robot team succeeds 0.91 to 0.92 of the time.

  5. The headline LLM number differs between tables

    The paper reports the same setting as 0.30 in one table and 0.25 in another. Scores on the test set are lower than on the validation set.216

    Details

    For ReAct with learned skills and ConceptGraphs perception on the validation set, Table 2 gives success 0.30±0.01 and Table 10 gives 0.25±0.01, with different step counts (12490.80 against 12274.27). The project page quotes 30%. Test-set results (Table 11) are lower than validation throughout: for example 0.51 against 0.70 for the fine-tuned model and 0.69 against 0.84 for the heuristic expert.

  6. Generated checks are sometimes wrong

    About 1 in 6 generated training episodes has a mistake in the task or in its automatic check.2

    Details

    Evaluation functions are written by an LLM. On 400 sampled generated episodes, 92% of evaluation functions and 83% of task and check pairs were correct; for spatial tasks only 74% were. Errors include wrong ordering constraints and wrong predicates. Validation and test episodes were annotated by people, but their residual error rate is not reported.

Details

About

What it is
Benchmark116
More

The paper and project page call PARTNR a benchmark.

Built by
Meta FAIR216
More

All 20 authors are marked 'Work done at FAIR Meta'; authors are listed alphabetically. The project page calls it 'A Meta FAIR Release'.

Released
October 2024, at ICLR 20251917+1
More

arXiv v1 and the Meta blog post on 2024-10-31. The episode dataset's 'initial release' commit is dated 2024-10-30.

Code repository created 2024-10-28 (GitHub API). arXiv has only v1.

Version
Episodes v0_0. The code has no tagged release.10188
More

Episode dataset v0_0, the only entry in its changelog. Code has no tagged releases; pull request 'PARTNR version v0.1.0' was merged on 2025-01-31.

setup.py gives version '1.0'.

Last update
July 2025. Task-type metadata was added to the dataset.111917
More

Episode dataset: task-type metadata added 2025-07-16. Code: last commit on main 2025-04-17.

The GitHub API's pushed_at of 2026-04-08 comes from automated dependency-update branches, not main.

Status
No updates since mid-2025 Inferred191112+1
More

No code changes on main since 2025-04-17; dataset metadata last changed 2025-07-16. Issues opened since 2025-09 have no maintainer reply.

It also depends on habitat-lab, which Meta stopped maintaining in 2026-05.

Setup

Runs in
Simulation2
More

Habitat 3.0 with HSSD scenes. Real people can control the human avatar in the same simulator.

Simulator
Habitat 3.0221
More

Habitat 3.0 (habitat-sim and habitat-lab), with HSSD scenes extended by articulated furniture.

The paper says the code depends on v0.3.2; the installation guide installs habitat-sim 0.3.3. Meta stopped maintaining habitat-lab after v0.3.4 (2026-05-07).

Robot
Arm on wheels2
More

Changed from the basic entry, which also listed 'humanoid'. The humanoid is a simulated person, not a robot under test.

Robot model
Spot (simulated)222
Setting
Whole home2
Tasks
100,000 tasks in the paper1210
More

100,000 tasks (abstract). Paper splits: 100,000 train episodes, 1,000 validation, 1,000 test. Released: 111,652 train episodes verified by people (131,991 unverified) and 1,000 validation. No test split is released.

CONFLICT between the paper's split sizes and the released files. The README lists only train_2k, val, train and val_mini as runnable splits.

Scenes
60 houses12
More

60 HSSD houses with added articulated furniture

Training data
Traces from 129 people solving tasks in the simulator21011
More

De-identified traces from 129 participants solving tasks in the simulator (released for validation and a 2,000-episode train subset), plus ReAct traces used for retrieval and fine-tuning.

Changes at test
New houses and new instructions Inferred2
More

Validation and test episodes use houses not in the training split, each with its own new instruction. Partners can be unseen real people.

100,000 train episodes in 37 scenes, 1,000 validation in 13, 1,000 test in 10 (paper Section 3.3).

Scoring and access

Scored by
Success rate, Progress score2
More

Success means all propositions are satisfied under their constraints. Percent complete is the share of propositions satisfied.

Score
Success and share of the task completed, checked by code written by an LLM2
More

Each episode has a Python evaluation function, generated by an LLM (CodeLlama-70B) from human-verified examples, that checks propositions, their order and constraints over the whole episode. It returns success, percent complete and a failure explanation. Papers also report simulation steps, planning cycles (capped at 50), task offloading, extraneous effort and exploration efficiency.

Manual check of 400 sampled generated episodes: 92% of evaluation functions and 90% of instructions correct, 83% both (Table 7). Validation and test episodes were annotated by people.

Trials
1 run on each of 1,000 validation episodes2
More

Main results · Table 2 and Table 10 give mean and standard error over the validation set (1,000 episodes). Table 11 gives the test set.2

Human-in-the-loop · Each task was attempted up to 3 times, with a written explanation of the failure after each try. The successful try, or else the best try, was kept.2

Who runs it
Each team tests its own model Inferred168
More

No organiser runs submissions.

Error bars
Usually reported Inferred234+1
More

The paper reports mean and standard error, MIT Lincoln Lab one standard deviation and AHAT ±. FLEET's table gives single numbers.

Leaderboard
Results are reported only in papers. Inferred16810+1
More

No leaderboard or challenge on the project page, README, dataset card or in the paper.

Code licence
MIT78
More

LICENSE: MIT, 'Copyright (c) Meta Platforms, Inc. and its affiliates'.

Data licence
CC BY-NC 4.010
More

Episodes, checkpoints, scene graphs and human traces: CC BY-NC 4.0 (dataset card).

Asset licence
HSSD scenes are non-commercial. The objects have no licence label.232425+2
More

HSSD scenes (partnr branch) CC BY-NC 4.0. OVMM objects have no card or licence file. Humanoid avatars non-commercial, with conflicting labels. Spot model by permission of Boston Dynamics.

OVMM objects

Humanoid avatars · Card metadata says CC BY-NC-SA 4.0; card text says CC BY-NC 4.0, with walking motion under the SMPL Body Motion File License.25

Access
Open download262110
More

Code on GitHub; data on Hugging Face without gating.

Running the paper's baselines needs Llama-3.1-70B; the human study hosted each 70B model on 4 A100 GPUs. The HSSD card shows a licence acknowledgement prompt.

Commercial use
Not allowed Inferred1023
More

Episodes and scenes are CC BY-NC 4.0. Not legal advice.

Published at
ICLR 20252728
More

ICLR 2025 (poster)

Sources 28

  1. 1PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks (arXiv abstract page)Paper · 31 Oct 2024 · checked 10 Oct 2026
  2. 2PARTNR paper, full text v1 (Sections 3 and 4; Appendices A.1, A.6, A.11, A.13)Paper · 31 Oct 2024 · checked 10 Oct 2026
  3. 3Evaluation of Habitat Robotics using Large Language Models (Table II)Paper · 8 Jul 2025 · checked 10 Oct 2026
  4. 4FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams (Table I, Section IV-D)Paper · 8 Oct 2025 · checked 10 Oct 2026
  5. 5Any House Any Task (v2 retitled TGPO), Table 1Paper · 12 Feb 2026 · checked 10 Oct 2026
  6. 6Semantic Scholar API: citations of arXiv:2411.00081 with contexts (98 records)Index · 10 Oct 2026 · checked 10 Oct 2026
  7. 7partnr-planner LICENSERepository · Oct 2024 · checked 10 Oct 2026
  8. 8facebookresearch/partnr-planner READMERepository · Mar 2025 · checked 10 Oct 2026
  9. 9Advancing embodied AI through progress in touch perception, dexterity, and human-robot interaction (PARTNR section)Official blog · 31 Oct 2024 · checked 10 Oct 2026
  10. 10ai-habitat/partnr_episodes dataset card (splits, licence, changelog)Dataset page · 16 Jul 2025 · checked 10 Oct 2026
  11. 11Hugging Face Hub API: ai-habitat/partnr_episodes file tree and commitsIndex · 10 Oct 2026 · checked 10 Oct 2026
  12. 12partnr-planner issue #38: fine-tuned planning model weights missingRepository · 22 Nov 2025 · checked 10 Oct 2026
  13. 13partnr-planner issue #35: evaluation metrics code missing for LLM-LLM runsRepository · 25 Sep 2025 · checked 10 Oct 2026
  14. 14partnr-planner issue #43: transformers-CFG does not support Qwen or Llama-3.2 and aboveRepository · 18 Mar 2026 · checked 10 Oct 2026
  15. 15partnr-planner issue #18: episodes already satisfied at the start (with co-author reply)Repository · 10 Mar 2025 · checked 10 Oct 2026
  16. 16PARTNR project pageOfficial site · 2024 · checked 10 Oct 2026
  17. 17GitHub API: facebookresearch/partnr-planner (stars, forks, created, pushed) and commit listIndex · 10 Oct 2026 · checked 10 Oct 2026
  18. 18partnr-planner pull request #6: PARTNR version v0.1.0 (merged)Repository · 31 Jan 2025 · checked 10 Oct 2026
  19. 19partnr-planner commit history (main)Repository · 17 Apr 2025 · checked 10 Oct 2026
  20. 20habitat-lab README (Meta maintenance notice beyond v0.3.4)Repository · 7 May 2026 · checked 10 Oct 2026
  21. 21partnr-planner INSTALLATION.md (habitat-sim 0.3.3, data downloads)Repository · Feb 2025 · checked 10 Oct 2026
  22. 22ai-habitat/hab_spot_arm dataset card (Spot URDF licence note)Dataset page · 14 Feb 2025 · checked 10 Oct 2026
  23. 23hssd/hssd-hab dataset card (CC BY-NC 4.0)Dataset page · 14 Feb 2025 · checked 10 Oct 2026
  24. 24Hugging Face Hub API: ai-habitat/OVMM_objects file tree (no card, no licence file)Index · 10 Oct 2026 · checked 10 Oct 2026
  25. 25ai-habitat/habitat_humanoids dataset cardDataset page · 18 Oct 2023 · checked 10 Oct 2026
  26. 26Hugging Face Hub API record for ai-habitat/partnr_episodes (downloads, gated flag)Index · 10 Oct 2026 · checked 10 Oct 2026
  27. 27PARTNR, ICLR 2025 proceedings pagePaper · 2025 · checked 10 Oct 2026
  28. 28ICLR 2025 poster page: PARTNRPaper · Apr 2025 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: PARTNR paper full text, project page, Meta blog (2024-10-31), dataset card, README, Semantic Scholar citation contexts (98 records) filtered for real-robot mentions, FLEET (only paper with hardware trials), web searches. None found.

leaderboard: Project page, README, dataset card, ICLR page, paper text.

used_by: Semantic Scholar citation contexts (98 records); opened 2507.06157, 2510.07417, 2602.12244, 2605.12920, 2606.28182. OpenAlex could not be queried (daily budget exhausted); arXiv API search returned errors.

license_assets (OVMM objects): ai-habitat/OVMM_objects repository tree: no card, no licence file.

issues (reviews): OpenReview forum T5QLRRHyL1; the API returned HTTP 403, so reviews were not read.

Change history

  1. Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/frontier-labs.json. Changes from the basic entry: embodiment no longer lists 'humanoid'; prior claim 'humans 0.93 vs LLM 0.30' confirmed with caveats (retries for people, 0.25 in Table 10, different partners); added test-set results, missing planner weights, the 'already satisfied' dispute, third-party results, issues and readings. Checks ran on 2026-10-10 and into early 2026-10-11 local time.
  2. Published as a full entry.