PARTNR
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
PARTNR is a set of 100,000 simulated household tasks in which a robot's planner (the model that decides its next steps) works with a person. It is mainly used to test task planning and coordination by large language models (LLMs).12
What a score here does not tell you Inferred
- How well the planner will work on a real robot.We found no runs on real robots.
- How good the robot is at low-level manipulation (moving and handling objects).The robot's skills are either oracle skills provided by the simulator or fixed learned modules.
- How a planner does on the test set.The test episodes have not been released.
Comparisons with real robots
Details
The paper has no real-robot runs. Meta's blog calls testing in physical-world scenarios a future goal. FLEET ran two real Spot robots on different inspection tasks, with no paired PARTNR comparison.
Our assessment Opinion
PARTNR tests planners in simulation. It does not test robot control.
Reasoning
PARTNR measures high-level planning and coordination with fixed skills in simulation. A high score says little about low-level manipulation or real robots. With learned skills and realistic perception, the best LLM success fell to 0.30 (or 0.25) in the paper.
Confidence: medium
People do better than LLM agents, but part of the gap comes from differences in how they were tested.
Reasoning
The gap between people and LLM agents is real, but the headline numbers overstate it. People had retries and feedback, and the automated runs pair the LLM robot with a weaker LLM partner.
Confidence: medium
Check the split, the skills and the metric definitions before you compare numbers across papers.
Reasoning
Do not compare PARTNR numbers across papers without checking the split, the skills, the perception setting and the metric definitions. The released data supports only validation-set comparisons.
Confidence: high
Known problems 6, 1 disputed
Released data does not match the paper
The release has no test split. The weights of the fine-tuned planner are missing.21011+3
Details
The paper describes 100,000 train, 1,000 validation and 1,000 test episodes. The release has 111,652 people-verified train episodes (131,991 unverified) and 1,000 validation episodes, and no test split, so the paper's Table 11 cannot be reproduced. The paper says weights of the fine-tuned planner were released; the dataset holds only the skill checkpoints and an issue reporting this has had no reply since 2025-11-22. Another open issue reports that the code for task offloading and related metrics covers only the human runs. The constrained decoding library does not support Qwen or Llama-3.2 and newer models.
Later papers change the metrics or the task set
Papers by other groups redefine the metrics or use only some of the tasks.345+1
Details
MIT Lincoln Lab redefines success rate as the share of successful decision points and percent complete as the share of successful episodes, ran each setting for ten hours (different episode counts per model), and reprints the paper's Llama-3.1-70B rows next to its own. FLEET reports three of the four task types. AHAT scores planning only. The README notes that constrained generation, which the paper uses to block invalid actions, is not supported for OpenAI models.
Some tasks are already done at the startDisputed
Some episodes are already solved at the start. The authors say this is intended.152
Other view: A co-author says these episodes are intentional and that there are few of them.15
Details
Users report validation episode 139, where the objects already met the goal and an agent succeeded by checking and calling Done. The paper's own error analysis lists 'Already Satisfied' as a task-generation failure (2% of sampled constraint-free, spatial and temporal episodes).
A PARTNR co-author replied that such episodes are intentional, to test whether agents can verify that a task is complete, and that filtering keeps their number limited. An impossible test episode was fixed.
People and LLMs are scored under different rules
People could retry each task and got feedback after each try. LLM agents got one run.216
Details
The headline 'humans solve 93%, LLMs 30%' compares different set-ups. In the human runs each task could be tried up to 3 times, with a written explanation of what went wrong after each try, and the best try was kept. Tasks that no person solved in 6 tries had already been removed from the dataset. The 0.30 comes from LLM planners with learned skills, scene graphs built from camera images and an LLM-controlled partner, in one run. When the partner is a real person, the same kind of LLM robot team succeeds 0.91 to 0.92 of the time.
The headline LLM number differs between tables
The paper reports the same setting as 0.30 in one table and 0.25 in another. Scores on the test set are lower than on the validation set.216
Details
For ReAct with learned skills and ConceptGraphs perception on the validation set, Table 2 gives success 0.30±0.01 and Table 10 gives 0.25±0.01, with different step counts (12490.80 against 12274.27). The project page quotes 30%. Test-set results (Table 11) are lower than validation throughout: for example 0.51 against 0.70 for the fine-tuned model and 0.69 against 0.84 for the heuristic expert.
Generated checks are sometimes wrong
About 1 in 6 generated training episodes has a mistake in the task or in its automatic check.2
Details
Evaluation functions are written by an LLM. On 400 sampled generated episodes, 92% of evaluation functions and 83% of task and check pairs were correct; for spatial tasks only 74% were. Errors include wrong ordering constraints and wrong predicates. Validation and test episodes were annotated by people, but their residual error rate is not reported.
Details
About
- Built by
- Meta FAIR216
More
All 20 authors are marked 'Work done at FAIR Meta'; authors are listed alphabetically. The project page calls it 'A Meta FAIR Release'.
- Released
- October 2024, at ICLR 20251917+1
More
arXiv v1 and the Meta blog post on 2024-10-31. The episode dataset's 'initial release' commit is dated 2024-10-30.
Code repository created 2024-10-28 (GitHub API). arXiv has only v1.
- Version
- Episodes v0_0. The code has no tagged release.10188
More
Episode dataset v0_0, the only entry in its changelog. Code has no tagged releases; pull request 'PARTNR version v0.1.0' was merged on 2025-01-31.
setup.py gives version '1.0'.
Setup
- Runs in
- Simulation2
More
Habitat 3.0 with HSSD scenes. Real people can control the human avatar in the same simulator.
- Simulator
- Habitat 3.0221
More
Habitat 3.0 (habitat-sim and habitat-lab), with HSSD scenes extended by articulated furniture.
The paper says the code depends on v0.3.2; the installation guide installs habitat-sim 0.3.3. Meta stopped maintaining habitat-lab after v0.3.4 (2026-05-07).
- Robot
- Arm on wheels2
More
Changed from the basic entry, which also listed 'humanoid'. The humanoid is a simulated person, not a robot under test.
- Setting
- Whole home2
- Tasks
- 100,000 tasks in the paper1210
More
100,000 tasks (abstract). Paper splits: 100,000 train episodes, 1,000 validation, 1,000 test. Released: 111,652 train episodes verified by people (131,991 unverified) and 1,000 validation. No test split is released.
CONFLICT between the paper's split sizes and the released files. The README lists only train_2k, val, train and val_mini as runnable splits.
- Training data
- Traces from 129 people solving tasks in the simulator21011
More
De-identified traces from 129 participants solving tasks in the simulator (released for validation and a 2,000-episode train subset), plus ReAct traces used for retrieval and fine-tuning.
- Changes at test
- New houses and new instructions Inferred2
More
Validation and test episodes use houses not in the training split, each with its own new instruction. Partners can be unseen real people.
100,000 train episodes in 37 scenes, 1,000 validation in 13, 1,000 test in 10 (paper Section 3.3).
Scoring and access
- Scored by
- Success rate, Progress score2
More
Success means all propositions are satisfied under their constraints. Percent complete is the share of propositions satisfied.
- Score
- Success and share of the task completed, checked by code written by an LLM2
More
Each episode has a Python evaluation function, generated by an LLM (CodeLlama-70B) from human-verified examples, that checks propositions, their order and constraints over the whole episode. It returns success, percent complete and a failure explanation. Papers also report simulation steps, planning cycles (capped at 50), task offloading, extraneous effort and exploration efficiency.
Manual check of 400 sampled generated episodes: 92% of evaluation functions and 90% of instructions correct, 83% both (Table 7). Validation and test episodes were annotated by people.
- Trials
- 1 run on each of 1,000 validation episodes2
More
Main results · Table 2 and Table 10 give mean and standard error over the validation set (1,000 episodes). Table 11 gives the test set.2
Human-in-the-loop · Each task was attempted up to 3 times, with a written explanation of the failure after each try. The successful try, or else the best try, was kept.2
- Error bars
- Usually reported Inferred234+1
More
The paper reports mean and standard error, MIT Lincoln Lab one standard deviation and AHAT ±. FLEET's table gives single numbers.
- Leaderboard
- Results are reported only in papers. Inferred16810+1
More
No leaderboard or challenge on the project page, README, dataset card or in the paper.
- Data licence
- CC BY-NC 4.010
More
Episodes, checkpoints, scene graphs and human traces: CC BY-NC 4.0 (dataset card).
- Asset licence
- HSSD scenes are non-commercial. The objects have no licence label.232425+2
More
HSSD scenes (partnr branch) CC BY-NC 4.0. OVMM objects have no card or licence file. Humanoid avatars non-commercial, with conflicting labels. Spot model by permission of Boston Dynamics.
OVMM objects
Humanoid avatars · Card metadata says CC BY-NC-SA 4.0; card text says CC BY-NC 4.0, with walking motion under the SMPL Body Motion File License.25
Sources 28
- 1PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks (arXiv abstract page)Paper · 31 Oct 2024 · checked 10 Oct 2026
- 2PARTNR paper, full text v1 (Sections 3 and 4; Appendices A.1, A.6, A.11, A.13)Paper · 31 Oct 2024 · checked 10 Oct 2026
- 3Evaluation of Habitat Robotics using Large Language Models (Table II)Paper · 8 Jul 2025 · checked 10 Oct 2026
- 4FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams (Table I, Section IV-D)Paper · 8 Oct 2025 · checked 10 Oct 2026
- 5Any House Any Task (v2 retitled TGPO), Table 1Paper · 12 Feb 2026 · checked 10 Oct 2026
- 6Semantic Scholar API: citations of arXiv:2411.00081 with contexts (98 records)Index · 10 Oct 2026 · checked 10 Oct 2026
- 7partnr-planner LICENSERepository · Oct 2024 · checked 10 Oct 2026
- 8facebookresearch/partnr-planner READMERepository · Mar 2025 · checked 10 Oct 2026
- 9Advancing embodied AI through progress in touch perception, dexterity, and human-robot interaction (PARTNR section)Official blog · 31 Oct 2024 · checked 10 Oct 2026
- 10ai-habitat/partnr_episodes dataset card (splits, licence, changelog)Dataset page · 16 Jul 2025 · checked 10 Oct 2026
- 11Hugging Face Hub API: ai-habitat/partnr_episodes file tree and commitsIndex · 10 Oct 2026 · checked 10 Oct 2026
- 12partnr-planner issue #38: fine-tuned planning model weights missingRepository · 22 Nov 2025 · checked 10 Oct 2026
- 13partnr-planner issue #35: evaluation metrics code missing for LLM-LLM runsRepository · 25 Sep 2025 · checked 10 Oct 2026
- 14partnr-planner issue #43: transformers-CFG does not support Qwen or Llama-3.2 and aboveRepository · 18 Mar 2026 · checked 10 Oct 2026
- 15partnr-planner issue #18: episodes already satisfied at the start (with co-author reply)Repository · 10 Mar 2025 · checked 10 Oct 2026
- 16PARTNR project pageOfficial site · 2024 · checked 10 Oct 2026
- 17GitHub API: facebookresearch/partnr-planner (stars, forks, created, pushed) and commit listIndex · 10 Oct 2026 · checked 10 Oct 2026
- 18partnr-planner pull request #6: PARTNR version v0.1.0 (merged)Repository · 31 Jan 2025 · checked 10 Oct 2026
- 19partnr-planner commit history (main)Repository · 17 Apr 2025 · checked 10 Oct 2026
- 20habitat-lab README (Meta maintenance notice beyond v0.3.4)Repository · 7 May 2026 · checked 10 Oct 2026
- 21partnr-planner INSTALLATION.md (habitat-sim 0.3.3, data downloads)Repository · Feb 2025 · checked 10 Oct 2026
- 22ai-habitat/hab_spot_arm dataset card (Spot URDF licence note)Dataset page · 14 Feb 2025 · checked 10 Oct 2026
- 23hssd/hssd-hab dataset card (CC BY-NC 4.0)Dataset page · 14 Feb 2025 · checked 10 Oct 2026
- 24Hugging Face Hub API: ai-habitat/OVMM_objects file tree (no card, no licence file)Index · 10 Oct 2026 · checked 10 Oct 2026
- 25ai-habitat/habitat_humanoids dataset cardDataset page · 18 Oct 2023 · checked 10 Oct 2026
- 26Hugging Face Hub API record for ai-habitat/partnr_episodes (downloads, gated flag)Index · 10 Oct 2026 · checked 10 Oct 2026
- 27PARTNR, ICLR 2025 proceedings pagePaper · 2025 · checked 10 Oct 2026
- 28ICLR 2025 poster page: PARTNRPaper · Apr 2025 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: PARTNR paper full text, project page, Meta blog (2024-10-31), dataset card, README, Semantic Scholar citation contexts (98 records) filtered for real-robot mentions, FLEET (only paper with hardware trials), web searches. None found.
leaderboard: Project page, README, dataset card, ICLR page, paper text.
used_by: Semantic Scholar citation contexts (98 records); opened 2507.06157, 2510.07417, 2602.12244, 2605.12920, 2606.28182. OpenAlex could not be queried (daily budget exhausted); arXiv API search returned errors.
license_assets (OVMM objects): ai-habitat/OVMM_objects repository tree: no card, no licence file.
issues (reviews): OpenReview forum T5QLRRHyL1; the API returned HTTP 403, so reviews were not read.
Change history
- Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/frontier-labs.json. Changes from the basic entry: embodiment no longer lists 'humanoid'; prior claim 'humans 0.93 vs LLM 0.30' confirmed with caveats (retries for people, 0.25 in Table 10, different partners); added test-set results, missing planner weights, the 'already satisfied' dispute, third-party results, issues and readings. Checks ran on 2026-10-10 and into early 2026-10-11 local time.
- Published as a full entry.