BEHAVIOR-1K
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
What a score here does not tell you Inferred
- How well a policy (the robot's control model) will do on a real robot.Only one task has been tested on a real robot, in 2022. The trained policy scored about 40% in simulation and 0% on the real robot.
- Whether the robot does the task safely or carefully.The score counts only the final states of the task's target objects.
- How a policy handles new tasks or objects.The tests reuse the training tasks with new start states.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| BEHAVIOR-1K paper, real-robot study Dec 2022 | The trained policy had about 40% success in simulation and 0% on the real robot. The 3 policies ranked in the same order. No statistic was reported.1718 | The authors tested one task (CollectTrash) in a scanned copy of a lab apartment. The trained policy ran 50 times in simulation and 26 times on a real Tiago++ robot. Three trained policies were compared by rank. The study’s authors described the result as “preliminary”. | The benchmark’s authors |
Benchmarks built on BEHAVIOR-1K: COHERENT, Behavior-Skill, ManiUnit.192021
Our assessment Opinion
A score shows progress on known household tasks in simulation. It says little about real homes.
Reasoning
A BEHAVIOR Challenge score shows how much of each known household task a simulated robot completes from new starting positions. Partial credit makes up a large part of the leading scores. It says little about a real home. The only real-robot test is one task from 2022, where the trained policy scored 0%.
Confidence: medium
Ignore small differences in Q-score between papers.
Reasoning
Do not read gaps of a few Q points (Q is the share of goal conditions met) between papers as real differences. Re-runs moved per-task scores by more than 0.27, and papers mix held-out, public and self-selected results from different software versions.
Confidence: high
Top scores are far below the maximum. The best verified entry fully completed about one task run in eight.
Reasoning
The benchmark is far from saturated, which means top scores are still far below the maximum. The best verified 2025 entry fully completed about one task run in eight (success 0.1240), so the benchmark can still separate strong systems.
Confidence: high
Commercial users should read the asset licence first.
Reasoning
For companies, the asset licence matters more than the MIT code licence. The tasks need the encrypted, non-commercial asset bundle.
Confidence: high
Known problems 7
The software changed between challenge editions
Scores from 2025 and 2026 come from different software versions and task sets.222324+3
Details
The 2025 challenge ran on BEHAVIOR-1K v3.7 with Isaac Sim 4.5.0; the 2026 challenge runs on v3.9 with Isaac Sim 5.1.0. A maintainer says the first 50 tasks of the 2026 instances are the 2025 instances replayed on the newer version for compatibility. Users reported an unstable scene that gave invalid partial credit (issue #2324) and divergent action replay of the 2026 demos (issue #2344); maintainers could not reproduce either and measured 0.026 m of drift in one replay. The 2026 dataset was corrected on 2026-07-27 and 2026-08-24. 2025 and 2026 scores also differ in task count (50 against 100) and in robot rules.
Demonstrations labelled MIT are rendered from non-commercial assets
The MIT licence on the demonstrations and the non-commercial licence on the assets may conflict. Inferred272823
Details
The challenge demonstrations are labelled MIT, but they are videos and states rendered from the encrypted asset bundle. That bundle's EULA allows only non-commercial academic research and forbids redistributing the data 'in whole or part'. The cards do not say how the two licences interact.
Papers compare numbers from different protocols
Later papers mix held-out scores from a hidden test set, public scores and scores that the authors selected themselves.8910+3
Details
The winner's held-out 0.2599 used four checkpoints and task-specific correction rules. Comet's 0.3453 is a post-challenge score on the validation instances, with the checkpoint chosen on those instances. Galaxea's G0.5 report compares its own two-run average (10 instances per task; the set is not stated) with leaderboard public scores, including Comet's 0.1830, which the Comet team says came from an incomplete evaluation. AHAT reports 70.3% 'success' from planning alone. The 2025 hidden instances were published on 2025-12-15, so later papers can no longer report a held-out score on the 2025 set.
Re-running the same policy gives different per-task scores
When another team re-ran the same checkpoints (saved versions of the trained policy), some task scores moved by more than 0.27.6148+2
Details
The organisers state that the simulator is nondeterministic. The Huawei team re-ran the winner's published checkpoints with the official scripts: per-task scores differed from the posted ones by more than 0.27 on tasks 24, 26 and 27, while the winner's average stayed close to its posted 0.2605. In their re-runs the winner averaged Q 0.256 and the runner-up 0.192, against 0.2514 for the runner-up on the held-out test. A Georgia Tech paper ran an openpi-comet checkpoint on all 50 tasks and got per-task values averaging 0.0788 by our arithmetic, but it may not have used the competition's checkpoint suite. The winner also reports that its cloud evaluation machines rendered without NGX, which visibly degraded images but had very small impact on success rate.
The tests use known tasks with new start states and allow hand-written rules
The tests reuse the training tasks. Hand-written rules raised the winning team's score.9127+1
Details
All 50 challenge tasks appear in both the training demonstrations and the test. The winning team calls the required generalisation very limited and notes there is no test of unseen object categories, language goals or new tasks. It also reports that training demonstrations were biased toward simpler instances. Its task-specific correction rules, such as reopening an accidentally closed gripper, raised Q 2.2 times on 13 tasks (39 episodes); a later paper declined to compare with it for that reason. The rules allow any method.
The score ignores how the goal was reached
Only the final states of objects count, so unsafe actions are not penalised.146
Details
Q counts only the final states of the task's target objects. A 2026 analysis by Huawei Noah's Ark Lab (Canada) points out that dropping a non-target object, such as a chopping board, or hitting furniture is not penalised. It proposed safety-adjusted scores and re-ran the top two 2025 entries: adjusting for target-object violations cut task scores by up to 35%, and up to 40% of cases showed some violation when non-target objects were added. Averages fell from Q 0.256 to 0.239 for the winner and from 0.192 to 0.173 for the runner-up.
Evaluation is slow and needs NVIDIA RTX hardware
Evaluating a single task on RTX GPUs can take hours.1069+1
Details
The Comet team reports that evaluating a single task can take from one hour to nearly a full day. On an RTX 4090, the organisers measured scene loading of about 150 to 300 seconds and 13.52 to 24.55 frames per second with random actions. The winner used an on-demand cluster and reports that a full evaluation finishes in under 2 days. Final 2026 evaluation runs on GPUs such as the RTX 3090, A5000 and TitanRTX.
Details
About
- What it is
- Benchmark118
More
The paper calls BEHAVIOR-1K a simulation benchmark with fixed activity definitions (BDDL) and metrics.
- Built by
- Stanford Vision and Learning Lab182
More
2024 paper: 35 authors. Funders listed include Stanford HAI, Toyota Research Institute, NSF, ONR, Amazon, Bosch, Salesforce and Samsung.
Stanford University (Stanford Vision and Learning Lab) · Lead organisation. Site footer: '© 2026 Stanford Vision and Learning Lab'. Most authors are at Stanford.182
Co-author affiliations (2024 paper) · The University of Texas at Austin, University of Illinois Urbana-Champaign, University of Southern California, Salesforce Research.18
NVIDIA Research (internships) · Acknowledgements: work done in part while three authors were interns at Nvidia Research.18
Challenge sponsors · 2026 page logos: Simovation, IMDA, Stanford HAI, Schmidt Family Foundation. NVIDIA joined as a 2025 sponsor (announcement of 2025-10-08).34
- Released
- December 2022, at CoRL 202231132+1
More
CoRL 2022 (Auckland, 14 to 18 December 2022), PMLR volume 205. Extended version on arXiv on 2024-03-14.
The CoRL title is 'BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation'. The GitHub repository was created on 2021-12-17. The first release tag, OmniGibson v0.0.1, is dated 2022-12-16. The 2025 challenge launched on 2025-09-02.
- Version
- v3.9.3-post2, released October 2026332322+2
More
BEHAVIOR-1K v3.9.3-post2. The 2025 challenge ran on the v3.7 series with Isaac Sim 4.5.0; the 2026 challenge runs on the v3.9 series with Isaac Sim 5.1.0.
v3.7.0 to v3.7.2 · 2025-09-02 to 2025-12-15. Used for the 2025 challenge. setup.sh installs Isaac Sim 4.5.0.3322
v3.9.0 to v3.9.3-post2 · 2026-07-03 to 2026-10-07. Used for the 2026 challenge. setup.sh installs Isaac Sim 5.1.0.3323
OmniGibson v0.0.1 to v1.1.1 · Earlier tags of the same repository, 2022-12-16 to 2024-10-04, released under the OmniGibson name.33
2025 BEHAVIOR Challenge · 50 tasks. Two tracks: standard (onboard RGB, depth, segmentation and proprioception) and privileged information. Robot fixed to R1 Pro.47
2026 BEHAVIOR Challenge · 100 tasks in 7 scenes (4 new). One track (RGB, depth and proprioception). Robot not fixed: R1 Pro by default or another OmniGibson robot.36
Setup
- Runs in
- Simulation1
- Simulator
- OmniGibson on Isaac Sim 5.1182322+2
More
OmniGibson, built on NVIDIA Omniverse and PhysX 5, installed with Isaac Sim 5.1.0 (v3.9 series) or 4.5.0 (v3.7 series). Simulates rigid and deformable bodies, cloth, fluids and object states such as temperature.
2024 paper: about 60 fps for a house scene with about 60 objects, ray-traced. Requirements page: Ubuntu 22.04+ or Windows 10+, 32 GB RAM, NVIDIA RTX 2070 or better with 8 GB VRAM. NVIDIA's documentation page for Isaac Sim 5.1.0 carries a banner saying that release is no longer supported (checked 2026-10-10).
- Robot
- Arm on wheels567
More
OmniGibson supports 12 robots: 4 mobile robots, 3 manipulators, 4 mobile manipulators and a bimanual proxy for VR teleoperation. The taxonomy value 'bimanual-arm' means two arms on a fixed base, so it is not used for the R1 Pro.
- Robot model
- R1 Pro (simulated)518
More
R1 Pro: holonomic base, 4-DOF torso, two 7-DOF arms and two parallel-jaw grippers. The 2022 real-robot study used a PAL Robotics Tiago++.
R1 Pro maker · Not named on the BEHAVIOR pages. The same team's BRS paper names the smaller R1 as a Galaxea robot, and Galaxea's G0.5 report fine-tunes on real R1-Pro robots. So the maker is most likely Galaxea.3611
- Setting
- Whole home, Office or lab, Mixed13
More
Gardens, restaurants and stores have no exact taxonomy value, so 'mixed' is added.
- Tasks
- 1,000 activities, with 50 or 100 of them in the challenge1183+3
More
1,000 activities defined in BDDL. The challenge uses 50 of them (2025) and 100 (2026).
The 1,000 are the 909 activities ranked highest in a survey of 1,461 people plus 91 activities from BEHAVIOR-100 (2024 paper). The 2026 set keeps the 50 tasks of 2025 and adds 50.
- Scenes
- 50 scenes18
More
50 interactive scenes with 373 rooms: 15 houses from BEHAVIOR-100 and 35 new scenes
- Training data
- 20,000 teleoperated (remote-controlled) demonstrations for 202633728+3
More
20,000 teleoperated demonstrations for the 2026 challenge (1,950 hours, 200 per task). The 2025 challenge had 10,000.
2026: 20,000 demos, 1,950 hours · 100 tasks. 210,916,774 frames. Average trajectory 351.54 s. 3.27 TB in LeRobot v3 format; raw HDF5 1.44 TB. Collected with the JoyLo whole-body teleoperation interface; data provided by Simovation.33728
2025: 10,000 demos, '1200+ hours' · 50 tasks, 200 per task. The dataset card lists 119,094,660 frames at 30 fps.427
2025 hours: conflict · 119,094,660 frames at 30 fps is about 1,103 hours by our arithmetic. The challenge page says 1200+ hours; Galaxea's G0.5 report says over 1,100 hours.27411
Dataset corrections in 2026 · Base velocity frame fixed (2026-07-27), depth videos fixed (2026-07-27), arm, gripper and trunk velocity fields fixed (2026-08-24).26
- Changes at test
- Only the start positions change679
More
Challenge test instances keep the same tasks as the training demonstrations. Only the starting states of task objects and the robot's starting pose change.
2026 rules: 'Each instance differs in terms of initial object states and initial robot poses.' The first-place team calls the required generalisation 'very limited' and notes there is no test of unseen object categories, language goals or new tasks. No official train and test split exists for the full 1,000 activities.
Scoring and access
- Scored by
- Progress score, Success rate678
More
Ranking uses the partial-credit Q-score. Full success rate is also reported.
- Score
- Share of goal conditions met (Q-score)6818+1
More
Task success score Q: the share of BDDL goal conditions satisfied at the end of an episode, using the best-matched goal clause, averaged over instances and tasks. Full success rate is reported too. Ties are broken by simulated time, base distance and hand movement, normalised by human averages from 200 demonstrations per task.
The 2022 and 2024 papers report success rate, Q and three efficiency metrics. In the first-place entry, partial successes make up roughly half of the score (team report).
- Trials
- 1 run on 10 instances per task786+3
More
2025 challenge · Teams ran 1 rollout on each of 10 public instances per task (500 rollouts) and reported the scores. Timeout: 2 times the average human completion time. The organisers re-ran the top 5 on 10 held-out instances per task.78
2026 challenge · 100 tasks x 10 instances x 1 rollout = 1,000 rollouts. Timeout: 1.5 times the mean human demonstration length. From 2026-10-09 each rollout must also average at least 1 FPS. The organisers evaluate top submissions on hidden instances.63026
Nondeterminism · The organisers state the simulator is nondeterministic, so rollouts of the same policy on one instance can differ. They forbid picking the best of repeated runs.6
2024 paper baselines · Trained with seeds 0, 1 and 2 and evaluated with seed 0.18
- Who runs it
- Teams report their own scores. The organisers re-run the top entries.83830
More
Scores on public instances are self-reported. Held-out scores for the top teams are run by the organisers.
2025 leaderboard: 'Public Validation' entries are self-reported; 'Held-out Test' entries are verified by the BEHAVIOR team. 2026: final evaluation by the organisers from a Docker image or a policy server reached over the internet.
- Error bars
- Sometimes reported Inferred1889+1
More
The original papers report means with a spread over training seeds. The challenge leaderboards and the team reports we read give single numbers. G0.5 averages two evaluation runs without an interval.
- Leaderboard
- Official, for the challenge task sets only838
More
Official boards for the challenge task sets only: 2025 (provisional, on behavior.stanford.edu) and 2026 (Hugging Face Space). No board covers the full 1,000 activities.
2026: self-reported scores were shown from 2026-08-31 (commit 'Show self-reported leaderboard scores') and withheld on 2026-10-09 pending verification, as announced on the updates page. Winners are due 2026-11-04. We do not reproduce the withheld 2026 scores here.
- Code licence
- MIT15
More
LICENSE: MIT, 'Copyright (c) 2023 Stanford Vision and Learning Group'.
- Data licence
- MIT (demonstrations and task instances)272829+1
More
Challenge demonstrations, task instances and the released 2025 hidden instances are labelled MIT on their Hugging Face cards. The 2026 raw HDF5 dataset has no card.
These files are rendered from, or point into, the encrypted asset bundle, which has its own licence (license_assets and issues.i6).
- Asset licence
- A custom end-user licence (EULA) for non-commercial use. The assets are encrypted.234018
More
Custom EULA, last revised 2022-12-08: non-commercial academic research only; the data is encrypted; use only inside OmniGibson; no reverse engineering; no redistribution of the key or of the data 'in whole or part'.
Asset origin · Objects were mainly bought from TurboSquid and edited for simulation. Users get an encrypted copy and need not buy the assets.4018
Simulator licence (dependency) · The isaacsim 5.1.0.0 package is labelled 'NVIDIA Proprietary Software'. NVIDIA's docs say the Isaac Sim source on GitHub is Apache 2.0, while the Kit SDK and NVIDIA 3D assets fall under the Isaac Sim Additional Software and Materials License (use on systems with NVIDIA GPUs; no redistribution or modification). setup.sh links the general NVIDIA Software License Agreement instead.414235+2
- Access
- Open, after accepting a click-through EULA322334+3
More
Code on GitHub. Demonstrations on ungated Hugging Face datasets. The encrypted asset bundle downloads after a click-through EULA in the installer. No account is needed.
setup.sh asks the user to accept the Conda terms, the NVIDIA Isaac Sim EULA and the BEHAVIOR Data Bundle EULA, or to pass flags that accept them.
Sources 44
- 1BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation (arXiv abstract page)Paper · 14 Mar 2024 · checked 10 Oct 2026
- 2BEHAVIOR project site, home pageOfficial site · 2026 · checked 10 Oct 2026
- 32026 BEHAVIOR Challenge, home pageOfficial site · Jul 2026 · checked 10 Oct 2026
- 42025 BEHAVIOR Challenge (archive page)Official site · Dec 2025 · checked 10 Oct 2026
- 5BEHAVIOR documentation: Robots (OmniGibson)Official site · 2026 · checked 10 Oct 2026
- 62026 BEHAVIOR Challenge: Evaluation and RulesOfficial site · 2026 · checked 10 Oct 2026
- 72025 BEHAVIOR Challenge: Evaluation and Rules (archive)Official site · 2025 · checked 10 Oct 2026
- 8Provisional 2025 BEHAVIOR Challenge leaderboardLeaderboard · 1 Dec 2025 · checked 10 Oct 2026
- 9Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge (v2)Paper · 7 Dec 2025 · checked 10 Oct 2026
- 10Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge (v3, 'Post-challenge bug fix')Paper · 5 Jan 2026 · checked 10 Oct 2026
- 11G0.5: One Autoregressive Stream for Robot Reasoning and Action (Section 5.3, Table 4)Paper · 12 Aug 2026 · checked 10 Oct 2026
- 12Make Your VLA More Robust Without More Data By Interleaving Motion Planning (MPVI; Appendix A, Table 1)Paper · 31 May 2026 · checked 10 Oct 2026
- 13Any House Any Task (v2 retitled TGPO: Trace-Guided Policy Optimization for Robot Task Planning), Table 1Paper · 12 Feb 2026 · checked 10 Oct 2026
- 14How VLAs (Really) Work In Open-World EnvironmentsPaper · 23 Apr 2026 · checked 10 Oct 2026
- 15BEHAVIOR-1K LICENSE fileRepository · 2023 · checked 10 Oct 2026
- 16Openpi Comet, v1 of the reportPaper · 10 Dec 2025 · checked 10 Oct 2026
- 17BEHAVIOR-1K, CoRL 2022 paper (PDF, Section 6.2)Paper · Dec 2022 · checked 10 Oct 2026
- 18BEHAVIOR-1K, extended paper, full text v1 (Sections 4 to 7, Table 1, Appendices F and G)Paper · 14 Mar 2024 · checked 10 Oct 2026
- 19COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language ModelsPaper · 23 Sep 2024 · checked 10 Oct 2026
- 20Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon TasksPaper · 31 Aug 2026 · checked 10 Oct 2026
- 21ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon TasksPaper · 8 Oct 2026 · checked 10 Oct 2026
- 22BEHAVIOR-1K setup.sh at tag v3.7.2 (Isaac Sim 4.5.0)Repository · 15 Dec 2025 · checked 10 Oct 2026
- 23BEHAVIOR-1K setup.sh on main (licence prompts, Isaac Sim 5.1.0 install)Repository · 27 Aug 2026 · checked 10 Oct 2026
- 24BEHAVIOR-1K issue #2324: invalid partial Q-score on putting_shoes_on_rack (Isaac Sim 5.1)Repository · 3 Aug 2026 · checked 10 Oct 2026
- 25BEHAVIOR-1K issue #2344: action-only replay of 2026 demos on the 3.9 stackRepository · 7 Sep 2026 · checked 10 Oct 2026
- 262026 BEHAVIOR Challenge: Announcements / UpdatesOfficial site · 9 Oct 2026 · checked 10 Oct 2026
- 27behavior-1k/2025-challenge-demos dataset card (meta/info.json)Dataset page · 2 Dec 2025 · checked 10 Oct 2026
- 28behavior-1k/2026-challenge-demos dataset cardDataset page · 5 Aug 2026 · checked 10 Oct 2026
- 29behavior-1k/2025-challenge-hidden-instances dataset (card and commit history)Dataset page · 15 Dec 2025 · checked 10 Oct 2026
- 302026 BEHAVIOR Challenge: Submission GuidelinesOfficial site · 2026 · checked 10 Oct 2026
- 31BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation (PMLR page)Paper · Dec 2022 · checked 10 Oct 2026
- 32GitHub API: StanfordVL/BEHAVIOR-1K (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 33StanfordVL/BEHAVIOR-1K releasesRepository · 7 Oct 2026 · checked 10 Oct 2026
- 34BEHAVIOR documentation: Installation (system requirements)Official site · 2026 · checked 10 Oct 2026
- 35Isaac Sim 5.1.0 documentation: NVIDIA Isaac Sim Additional Software and Materials LicenseOfficial site · 2026 · checked 10 Oct 2026
- 36BEHAVIOR Robot Suite (BRS): Streamlining Real-World Whole-Body Manipulation for Everyday Household ActivitiesPaper · 7 Mar 2025 · checked 10 Oct 2026
- 372026 BEHAVIOR Challenge: DatasetOfficial site · 2026 · checked 10 Oct 2026
- 38BEHAVIOR-1K 2026 Challenge leaderboard Space (README, data files, commit log)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 39behavior-1k/2026-challenge-rawdata README (returns 'Entry not found': no card)Dataset page · 22 Jun 2026 · checked 10 Oct 2026
- 40BEHAVIOR documentation: Asset SourcesOfficial site · 2026 · checked 10 Oct 2026
- 41PyPI JSON metadata for isaacsim 5.1.0.0 (licence field)Index · 10 Oct 2026 · checked 10 Oct 2026
- 42Isaac Sim 5.1.0 documentation: NVIDIA Isaac Sim LicensingOfficial site · 2026 · checked 10 Oct 2026
- 43NVIDIA Software License Agreement (page linked from setup.sh)Official site · unknown · checked 10 Oct 2026
- 44Hugging Face Hub API listing of behavior-1k datasets (2026 demos and raw data, created, modified, downloads)Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: CoRL 2022 paper (Section 6.2) and 2024 arXiv version (Section 6.2, Appendix G); BEHAVIOR site including Related Research (BRS, ACDC, MoMaGen, BEHAVIOR Vision Suite); 2025 and 2026 challenge pages; top-2 team reports (2512.06951, 2512.10071); G0.5 (2608.11739; real-robot results are on separate R1-Lite/R1-Pro tasks, not paired with BEHAVIOR scores); MPVI (2606.00985); Huawei analysis (2604.21192); Semantic Scholar citation contexts for both BEHAVIOR-1K records (650 citing papers) filtered for real-world mentions; web searches. No paired sim-vs-real study beyond the authors' 2022 one.
challenge results report by the organisers: Challenge pages, Stanford HAI news (2025-09-22, a launch story), web search for a 2025 results or lessons paper. None found.
leaderboard (2026): Hugging Face Space files data/results.jsonl (empty), data/self_reported_results.jsonl and the commit log; updates page. Verified 2026 results do not exist yet.
license_assets: setup.sh EULA text, Asset Sources page, 2024 paper Appendix D, Hugging Face cards, PyPI isaacsim metadata, Isaac Sim 5.1.0 licence pages, the NVIDIA Software License Agreement page linked by setup.sh.
robots (R1 Pro maker): Robots page, challenge pages and 2024 paper: maker not named. BRS paper and G0.5 report used for the inference.
Change history
- Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/core-sim-b.json. Folded the 2025 and 2026 BEHAVIOR Challenge into this record. Changes from the basic entry: embodiment narrowed to mobile-manipulator (bimanual-arm means a fixed base); sim_to_real level set to inferred with the full study design; '22% vs 40%' prior claim corrected (two different policies); added the 2025 hours conflict, Isaac Sim versions per edition, Semantic Scholar count for the CoRL record (459), validity v1, issues and readings. Checks ran on 2026-10-10 and into early 2026-10-11 local time.
- Published as a full entry.