LIBERO
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
130 simulated tasks for one robot arm. Many papers use it to report results for vision-language-action (VLA) models.123
What a score here does not tell you Inferred
- How well a policy will do on a real robot.Only one study has compared LIBERO scores with real-robot results. It found a correlation of r = 0.66 to 0.70 over five policies, and its authors call this poor.
- How well a policy handles new objects or new layouts.The test tasks are the same tasks the policy was trained on. Only the start positions change.
- Whether a policy understands the instruction.Studies show that policies can identify the task from the scene or from a task number, without reading the instruction.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| PolaRiS Dec 2025 | Correlation r = 0.66 to 0.704 | The same 5 policies were scored on LIBERO-90 and on real robots. The real-robot tasks were different from the LIBERO tasks. The study’s authors described the result as “poor”. | An independent group |
Benchmarks built on LIBERO: LIBERO-PRO, LIBERO-Plus, LIBERO-X, LangGap, LIBERO-Para, Libero-V, LIBERO-Safety, LIBERO-RECOVER.202122+5
Our assessment Opinion
A high score shows that a policy learned these simulated tasks. It says little about how the policy will do on real robots.
Reasoning
A high LIBERO score shows that a policy can learn these simulated tasks from their demonstrations. It does not show general manipulation skill. It is weak evidence about real-robot performance, because the only paired study found a correlation of r = 0.66 to 0.70 over five policies.
Confidence: medium
Do not rank models by differences of one or two points at the top.
Reasoning
Do not rank models by differences of one or two points. At the top of the leaderboard, such differences are often smaller than the variation between random seeds and the differences between the test setups that papers use.
Confidence: high
Read the scores on the harder variants together with the standard score.
Reasoning
The harder variants, such as LIBERO-PRO and LIBERO-Plus, show failures that the standard score does not show. Read their scores together with the standard score. None of the variants has been compared with real-robot results, and models can be tuned to them as well.
Confidence: medium
Most papers now report success on 40 known tasks. Check which protocol a number comes from.
Reasoning
LIBERO was designed to measure how a policy carries knowledge across a stream of new tasks (lifelong learning). In the papers we checked, the headline numbers measure success on 40 known tasks instead. Check which protocol a number comes from.
Confidence: high
LIBERO is still useful for quick, low-cost testing during development.
Reasoning
LIBERO is still useful for fast iteration, debugging and comparison with your own earlier runs. It is open, cheap to run and widely reproduced. The authors of the 2026 audit say the same about the benchmarks they audited.
Confidence: medium
Known problems 11, 1 disputed
Top scores are close to 100%
Since late 2025, the top average scores have been 98% to 99%. The best scores on single suites reach 99.5% to 100%.131415+6
Details
At least five papers since October 2025 report four-suite averages of 98.1 to 99.3%. Per-suite bests are 99.5 to 100% on Spatial, Object and Goal, and 98.8% on Long. Several later papers call LIBERO saturated.
Papers test in different ways
Papers run 10, 20 or 50 trials per task, train on different data and average over different sets of suites.72930+7
Details
Papers differ in trials per task (10, 20 or 50), training data (original or the filtered 1,693 episodes), one policy per suite vs one for all suites vs one expert per task, averages over three or four suites, and RL fine-tuning inside the test simulator. There is no validation split, so picking the best checkpoint on the test suite is allowed; for its own probe, the audit found this adds 0.6 to 4.4 points. Smaller settings matter too: the openpi script notes the environment seed moves objects even with fixed initial states; LeRobot says soft and hard resets give slightly different results and asks authors to pin the dataset revision; SimVLA finds single training choices can outweigh architecture changes.
A small model given only a task number scores close to the best
A model with 0.09 billion parameters was given only a task number and no instruction. It scored within about one point of the best models on three suites.3
Details
A 0.09B probe (DINOv2 encoder + MLP) gets a task ID instead of the instruction. It scored 99.0 / 100.0 / 98.8 / 92.4 on Spatial / Object / Goal / Long, within about one point of the best published result on three suites. It was trained and tuned per suite, and these are the best of several checkpoints scored on the test suite. With no checkpoint selection, its mean was 95.1%. The audit says this shows a high LIBERO score is not on its own evidence of general skill. It does not claim high-scoring policies lack skill.
Scores drop sharply after small camera or start-pose changes
In LIBERO-Plus, success falls from about 95% to below 30% after small changes to the camera or to the robot's start pose.2133
Details
LIBERO-Plus (CVPR 2026) perturbs seven factors. In its single-factor analysis, success drops from about 95% to below 30% under modest camera-viewpoint or robot-initial-state changes. Its released 10,030-task test set dropped tasks that all or most baseline models solved, so its scores are not on the same footing as standard LIBERO scores.
Most claimed improvements are not shown to be statistically significant
Only 19.8% of 789 claimed improvements on LIBERO can be shown to be statistically significant from the published numbers.3
Details
Of 789 previous-best-to-new comparisons on LIBERO (Spatial, Object and Goal pooled), 19.8% are provably significant at the 5% level from public scores. The rest show no improvement, are provably not significant, or cannot be tested because only averages are published. Each comparison uses the previous best as stated in the new paper.
Results change on different computers
In a test of 10 tasks, changing only the computer's CPU changed the simulation results in all 10.3
Details
With policy, seed and initial state fixed, changing only the CPU made the simulator state diverge in 10 of 10 LIBERO tasks for OpenVLA-OFT; in 5 of 10 the difference reached images and actions. Changing only the GPU also caused divergence. The audit argues bitwise determinism should not be the only reproducibility test, and finds aggregate moves small.
The scene often shows which task to do
Each layout is used for only one task, so a policy can tell the task from the scene. Sources disagree about whether this is true for the Goal suite.32324+1
Details
LIBERO draws instructions from a fixed set, one per task, so a task ID can replace language (issues.i2). LangGap says LIBERO assigns only one task per layout. LIBERO-Para says all LIBERO-Goal tasks start from one initial state, so there the instruction is the only cue. The two claims conflict for Goal, which the LIBERO paper describes as same objects and layout with different goals.
Models rely on the exact wording of the instruction
Without the instruction, success on LIBERO-Goal falls to 0% to 10%. When the instruction is reworded, success falls by 22 to 52 points.212024
Details
LIBERO-Plus's text says models largely ignore instructions. Its own blank-instruction test shows success falls sharply for most models when the instruction is removed, to about 0 to 10% on LIBERO-Goal for all six models; only OpenVLA-OFT on Object is unchanged. When the target object in the instruction is swapped, success drops to near zero. LIBERO-PRO finds near-identical trajectories under nonsense instructions. LIBERO-Para (EMNLP 2026) finds paraphrases cut success by 22 to 52 points across seven VLA set-ups, mostly through object synonyms; its authors read this as surface-level matching that disrupts task identification.
High-scoring models fail when object positions or tasks changeDisputed
In LIBERO-PRO, models that score above 90% on LIBERO drop to 0% when object positions or tasks change. The LIBERO-PRO authors say the models memorised the tasks.2016
Other view: The 2026 audit agrees that scores drop, but it describes the cause as weak generalisation. When the audit resampled start positions within the training range, scores changed by less than one point.3
Details
LIBERO-PRO perturbs objects, positions, instructions and tasks. Its abstract says models above 90% fall to 0.0% in its generalised setting. The drop depends on the perturbation: position and task changes push success to near zero, while object and instruction-meaning changes barely lower it. OpenVLA and π0 fail once an object moves more than 0.2 units; π0.5 holds to about 0.4 units and keeps 0.38 on LIBERO-Goal under position change. Nonsense instructions leave trajectories nearly unchanged. The authors read this as rote memorisation.
The 2026 audit accepts the drops but disputes the label. It says LIBERO-PRO and LIBERO-Plus change inputs outside the training distribution, so a drop may show weak generalisation rather than overfitting. When it redrew initial states inside the distribution (10,000 rollouts per policy), success moved by under one point (main text: Spatial Forcing −0.62, SimVLA +0.30, LeRobot π0.5 −0.14; its appendix table prints the opposite signs). It notes rollout noise leaves each sign uncertain.
The same model gets different scores in different papers
Different papers report the score of the same model, π0, as 86.0, 94.15 and 96.8.81030+8
Details
π0: 94.15 (openpi, 50 trials per task by script default), 86.0 (SmolVLA paper, 10 trials per task), 96.8 (T-MEE paper's own run). π0.5: 96.85 (openpi), 97.5 (LeRobot, 10 episodes per task), 97.7 (MoH paper's run). OpenVLA-OFT: 97.1 (its paper) vs 91.0 (SimpleVLA-RL's re-implementation with one camera and no proprioception; πRL reprints it labelled only 'OpenVLA-OFT'). SimVLA's own text and Table 2 disagree on two suites.
The included 3D models have no stated licence
The 3D models that come with LIBERO have no licence file. Some of them appear to come from collections with restricted licences. Inferred343536+2
Details
The repo and Hugging Face's re-hosted copy ship 3D assets with no licence file. Folder names suggest third-party origins. stable_hope_objects holds 14 grocery items; NVIDIA's HOPE set is 28 toy grocery items with 3D meshes under CC BY-NC-SA 4.0. turbosquid_objects holds 17 models; TurboSquid's Royalty Free License bars giving away model files outside a permitted Creation. stable_scanned_objects (11 items) has unchecked origin.
Details
About
- What it is
- Benchmark1
More
The paper calls it a benchmark with fixed task suites, demonstrations and metrics.
- Built by
- The University of Texas at Austin, Sony AI, Tsinghua University13839
More
Authors: Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, Peter Stone. The project site's Research page omits Yihao Feng; arXiv, NeurIPS and the README list all seven.
The University of Texas at Austin · LARG, RPL, and Statistical Learning & AI groups (project site footer). Six of seven authors.138
Sony AI · Second affiliation of Peter Stone.1
Tsinghua University · Affiliation of Chongkai Gao.1
- Released
- June 2023, at NeurIPS 20234041
More
arXiv v1 on 2023-06-05. Published at NeurIPS 2023, Datasets and Benchmarks Track.
The GitHub repository was created on 2023-04-08, before the paper. arXiv v2 is dated 2023-10-14.
- Version
- 0.1.0. There are no tagged releases.424344
More
Upstream package version 0.1.0. No tagged releases. Four suites: LIBERO-Spatial, LIBERO-Object, LIBERO-Goal (10 tasks each) and LIBERO-100 (split into LIBERO-90 and LIBERO-10).
setup.py and the docs both say 0.1.0. The GitHub tags and releases lists are empty. Two other upstream branches exist: 'docs' (last commit 2023-10-13) and 'X-embodiment' (see items). The packages on PyPI are Hugging Face builds, not upstream (see items).
LIBERO-Spatial · 10 tasks. Same objects; two identical bowls in different places. Varies spatial layout. Task: put the named bowl on the plate.17
LIBERO-Object · 10 tasks. Same layout; a different object to pick and place in each task. Varies object type.17
LIBERO-Goal · 10 tasks. Same objects, same layout. Varies the goal (the motion or behaviour asked for).17
LIBERO-100 · 100 tasks with mixed objects, layouts and backgrounds. Split into LIBERO-90 (90 short tasks, used for pretraining) and LIBERO-10 (10 long tasks, used for evaluation).145
LIBERO-10 = LIBERO-Long · Two names for the same 10 long-horizon tasks. The paper says LIBERO-Long; the code says libero_10.1467
Branch X-embodiment (unmerged) · Cross-embodiment code (first commit 2024-08-22) and a robosuite 1.5 update (2025-03-15). Never merged into master.47
hf-libero 0.1.4 (Hugging Face fork) · PyPI hf-libero, latest 0.1.4 on 2026-06-10, built from huggingface/LIBERO, a GitHub fork of upstream. LeRobot's 'libero' extra requires it. The fork downloads 3D assets at run time from the Hugging Face dataset lerobot/libero-assets. PyPI 'libero' 0.1.1 (2025-11-03) points to huggingface/lerobot-libero, now archived.484950+2
- Last update
- May 2025. One data file was replaced.525344
More
2025-05-18: one LIBERO-90 demonstration file replaced on the official Hugging Face mirror, with no changelog. Last code commit: 2025-03-15 (Hugging Face download support).
Replaced file: libero_90/LIVING_ROOM_SCENE1_pick_up_the_tomato_sauce_and_put_it_in_the_basket_demo.hdf5. Commit message: the original file 'seems to be corrupted'. Size 590,254,080 bytes before, 806,284,784 after (Hub tree at both revisions). Copies made before that date may hold the old file. The mirror was created 2025-03-13; its card was last edited 2025-03-17.
- Status
- No code changes since March 2025. Still widely used. Inferred445452+5
More
Upstream code unchanged since 2025-03-15; last data fix 2025-05-18. A Hugging Face fork is maintained. Use is very active.
Last merged pull request 2025-01-03. 108 open items (24 PRs, 84 issues); community PRs from October 2026 are unmerged. Last comment by a repo collaborator: 2025-08-06 (issue #99, Python >= 3.9 and new GPUs). requirements.txt pins robosuite 1.4.0 (released 2022-12-01); robosuite on PyPI is 1.5.2. PR #146 for robosuite 1.5.2 / MuJoCo 3.x is open. The hf-libero fork (0.1.4, 2026-06-10) keeps LIBERO installable for LeRobot users. For current use, see facts.used_by.
Setup
- Runs in
- Simulation1
- Simulator
- robosuite 1.4, built on MuJoCo1552
More
Paper: built on robosuite. requirements.txt pins robosuite==1.4.0. Environment code renders with MuJoCo. Tasks are written as BDDL/PDDL files.
- Robot
- One arm2
- Robot model
- Franka Emika Panda (simulated)25911
More
Code default is robots=['Panda']. Tabletop and kitchen problem classes wrap it as MountedPanda; the floor problem (LIBERO-Object) wraps it as OnTheGroundPanda. The paper text does not name the robot; OpenVLA-OFT names it.
- Setting
- Tabletop, Kitchen, Whole home Inferred134
More
Living room and study are mapped to 'home'. These are single workspaces, not whole houses.
- Tasks
- 130 tasks in 4 suites40134
More
130 tasks in four suites (10 + 10 + 10 + 100)
Repo has 10 + 10 + 10 + 90 + 10 BDDL task files (counted 2026-10-10). Most VLA papers use only the 40 tasks of Spatial, Object, Goal and Long.
- Scenes
- 20 layouts in LIBERO-100 Inferred153
More
20 scene layouts in LIBERO-100: 10 kitchen, 6 living room, 4 study
Counted by us from the scene names in the paper's LIBERO-90 task table (Appendix E.3) and the dataset file names. LIBERO-10 reuses scenes from that set. Spatial, Object and Goal use their own layouts and are not counted here.
- Training data
- 50 human demonstrations per task Inferred145
More
50 human demonstrations per task; 6,500 in total by our arithmetic. The project site says 65,000.
CONFLICT. The paper states 50 trajectories per task, collected by human teleoperation with a 3Dconnexion SpaceMouse. 130 x 50 = 6,500. The project Datasets page states 65,000. OpenVLA also counts 500 demonstrations per 10-task suite, which supports 50 per task.
50 per task · Human teleoperation with a 3Dconnexion SpaceMouse1
Observation content · Workspace and wrist RGB cameras, proprioception, language instruction, PDDL scene description45
128 x 128 px images · Original image resolution. OpenVLA re-rendered all demonstrations at 256 x 256.7
1,693 episodes (filtered) · OpenVLA replayed demonstrations and dropped failures: 68, 46, 72 and 121 of 500 in Spatial, Object, Goal and Long. The widely used Physical Intelligence copy has 1,693 episodes, 273,465 frames at 10 fps, over 40 tasks.76019
about 100.4 GB · Total size of the 130 HDF5 files on the official Hugging Face mirror (decimal GB, 10^9 bytes)53
- Changes at test
- Only the start positions change Inferred46720
More
In the common VLA protocol, test tasks are the training tasks. Only the initial object placement differs, drawn from a fixed set per task.
README: initial states are fixed for benchmarking. OpenVLA keeps 'the same initial environment configurations' as the original benchmark. In the original lifelong protocol, suites isolate shifts in layout, object and goal between tasks, but each new task comes with its own demonstrations, so this is transfer, not zero-shot generalisation.
Scoring and access
- Scored by
- Success rate17
More
Success is a binary goal predicate checked by the simulator. The original lifelong metrics (FWT, NBT, AUC) are built from success rates and have no taxonomy value of their own.
- Score
- Success rate for each suite, averaged across suites1712
More
Original paper: three lifelong-learning metrics, forward transfer (FWT), negative backward transfer (NBT) and area under the success curve (AUC), computed while tasks arrive one by one. Later VLA papers: train on the demonstrations, then report the success rate per suite and the plain average over Spatial, Object, Goal and Long.
The two protocols answer different questions. The first asks how well a learner keeps and transfers skills over a sequence. The second asks how well one model fits 40 known tasks. Almost all headline numbers today use the second.
- Trials
- 10 to 50 per task, depending on the paper1710+5
More
Fixed initial states: 50 per task · The benchmark loads a fixed set of initial states per task. This caps the common protocol at 50 distinct starts per task, 500 per suite.636219+1
Original paper · 20 rollouts per task, max 600 steps, every 5 epochs; 3 seeds (100, 200, 300)1
OpenVLA convention · 500 trials per suite (50 per task) from the fixed initial states; averaged over 3 seeds. Third-person images rotated 180 degrees at train and test time, because the environments rendered upside down on their hardware.7
openpi evaluation script · 50 trials per task, seed 7, images resized to 224 px, step limits 220 to 520 by suite. A code comment warns the environment seed changes object positions even with a fixed initial state.10
LeRobot · Docs recommend 10 episodes per task (400 in total) and averaging over 3 seeds. The code default is 50 episodes (EvalConfig.n_episodes).2961
SmolVLA paper · 10 trials per task30
NVIDIA GR00T N1.7 example · 200 episodes per suite (20 per task)31
- Who runs it
- Each team tests its own model Inferred3846
More
No organiser runs submissions. Each paper runs its own evaluation.
- Error bars
- Sometimes reported Inferred173+2
More
The original paper reports mean and standard error over 3 seeds. OpenVLA reports standard error. The 2026 audit finds most benchmarks report only an aggregate success rate. Since 2026-09-16, LeRobot's evaluator prints a 95% Wilson interval and success counts next to every success rate.
- Leaderboard
- None. Scores are only in papers. Inferred384665
More
No leaderboard on the project site or in the repo. The former Papers with Code LIBERO page now redirects to Hugging Face Trending Papers (HTTP 302, checked 2026-10-10).
- Code licence
- MIT6
More
LICENSE file: MIT, copyright 2023 Lifelong Robot Learning.
- Data licence
- Sources conflict: CC BY 4.0 or Apache-2.04653
More
Conflict: the README says CC BY 4.0; the official Hugging Face dataset card says Apache-2.0. The Hugging Face copy is now the only live official download.
CONFLICT between two primary sources from the same team. The README links the Hugging Face mirror as an official download, and download_utils.py uses it. Derived copies by other groups carry their own labels (items); they do not decide LIBERO's own licence.
MIT · openvla/modified_libero_rlds (OpenVLA's filtered copy)66
CC-BY-4.0 · physical-intelligence/libero (1,693 episodes)60
Apache-2.0 · lerobot/libero (recommended by LeRobot docs)67
Apache-2.0 · HuggingFaceVLA/libero68
- Asset licence
- Unknown
More
The repo ships 3D assets in folders named turbosquid_objects, stable_hope_objects and stable_scanned_objects. The full repo tree has no licence file besides the root MIT LICENSE, which covers 'the Software'. In the paper's NeurIPS checklist, 4(a) 'did you cite the creators' is Yes, 4(b) on asset licences is N/A; the paper text names no asset source. Hugging Face's re-hosted asset copy (lerobot/libero-assets) has no card and no licence. See issues.i11 for likely origins.
- Access
- Open. The data is on Hugging Face.534645+1
More
Code on GitHub. Data only from the ungated Hugging Face dataset yifengzhu-hf/LIBERO-datasets. The four original UT Box links (project Datasets page and download_utils.py) return HTTP 404 (checked 2026-10-10). The Datasets page's four 'Best Model Checkpoints' links are empty, so no official checkpoints can be downloaded. No registration needed.
- Commercial use
- Unclear Inferred64653+2
More
Code (MIT) and data (CC BY 4.0 or Apache-2.0, whichever applies) both allow commercial use with attribution. The bundled third-party 3D assets have no stated licence, and some folder names point to sources with non-commercial or no-redistribution terms (issues.i11). So the whole package is unclear. Not legal advice.
Sources 69
- 1LIBERO paper, full text v2Paper · Oct 2023 · checked 10 Oct 2026
- 2LIBERO environment wrapper source (robot and renderer)Repository · 2023 · checked 10 Oct 2026
- 3What Are We Actually Benchmarking in Robot Manipulation?Paper · Jun 2026 · checked 10 Oct 2026
- 4PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Section 5.2, Figures 6 and 7, Appendix C.1)Paper · Dec 2025 · checked 10 Oct 2026
- 5VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of VLA ModelsPaper · May 2026 · checked 10 Oct 2026
- 6LIBERO LICENSE fileRepository · Jun 2023 · checked 10 Oct 2026
- 7OpenVLA: An Open-Source Vision-Language-Action Model (Appendix E, LIBERO)Paper · Jun 2024 · checked 10 Oct 2026
- 8openpi LIBERO example README at first commit (π0 and π0-FAST results)Repository · 4 Feb 2025 · checked 10 Oct 2026
- 9openpi training configs (pi0_libero, pi05_libero use physical-intelligence/libero)Repository · 2025 · checked 10 Oct 2026
- 10openpi LIBERO evaluation scriptRepository · 2025 · checked 10 Oct 2026
- 11Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)Paper · Feb 2025 · checked 10 Oct 2026
- 12openpi LIBERO example README (current: π0.5 results)Repository · Sep 2025 · checked 10 Oct 2026
- 13X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelPaper · Oct 2025 · checked 10 Oct 2026
- 14Mixture of Horizons in Action Chunking (MoH)Paper · Nov 2025 · checked 10 Oct 2026
- 15Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time ExecutionPaper · Feb 2026 · checked 10 Oct 2026
- 16SimVLA: A Simple VLA Baseline for Robotic ManipulationPaper · Feb 2026 · checked 10 Oct 2026
- 17CORAL: Scalable Multi-Task Robot Learning via LoRA ExpertsPaper · Mar 2026 · checked 10 Oct 2026
- 18SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningPaper · Sep 2025 · checked 10 Oct 2026
- 19πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action ModelsPaper · Oct 2025 · checked 10 Oct 2026
- 20LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (v2, 2026-05-25)Paper · Oct 2025 · checked 10 Oct 2026
- 21LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action ModelsPaper · Oct 2025 · checked 10 Oct 2026
- 22LIBERO-X: Robustness Litmus for Vision-Language-Action ModelsPaper · Feb 2026 · checked 10 Oct 2026
- 23LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action ModelsPaper · Feb 2026 · checked 10 Oct 2026
- 24LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA ModelsPaper · Mar 2026 · checked 10 Oct 2026
- 25VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial ModelingPaper · Jun 2026 · checked 10 Oct 2026
- 26LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action ModelsPaper · Jun 2026 · checked 10 Oct 2026
- 27LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation ModelsPaper · Sep 2026 · checked 10 Oct 2026
- 28Reshaping Action Error Distributions for Reliable Vision-Language-Action Models (T-MEE)Paper · Feb 2026 · checked 10 Oct 2026
- 29LeRobot documentation: LIBERORepository · 2026 · checked 10 Oct 2026
- 30SmolVLA: A vision-language-action model for affordable and efficient roboticsPaper · Jun 2025 · checked 10 Oct 2026
- 31Isaac-GR00T LIBERO example and results (GR00T N1.7)Repository · Jul 2026 · checked 10 Oct 2026
- 32HarnessPAI: An Evolving Harness for Physical AI (three-suite average)Paper · Sep 2026 · checked 10 Oct 2026
- 33LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models (CVPR 2026 poster page)Paper · Jun 2026 · checked 10 Oct 2026
- 34LIBERO repo tree: BDDL task files and asset foldersRepository · Mar 2025 · checked 10 Oct 2026
- 35lerobot/libero-assets (Hugging Face dataset; no card)Repository · Nov 2025 · checked 10 Oct 2026
- 36NVIDIA HOPE dataset README (licence section)Repository · 2021 · checked 10 Oct 2026
- 37TurboSquid Royalty Free LicenseOfficial site · unknown · checked 10 Oct 2026
- 38LIBERO project site, main pageOfficial site · 2023 · checked 10 Oct 2026
- 39LIBERO project site, Research pageOfficial site · 2023 · checked 10 Oct 2026
- 40LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning (arXiv abstract page)Paper · Jun 2023 · checked 10 Oct 2026
- 41LIBERO, NeurIPS 2023 proceedings pagePaper · Dec 2023 · checked 10 Oct 2026
- 42LIBERO setup.pyRepository · 2023 · checked 10 Oct 2026
- 43LIBERO documentation: DatasetsOfficial site · 2023 · checked 10 Oct 2026
- 44LIBERO commit history, tags and releasesRepository · 15 Mar 2025 · checked 10 Oct 2026
- 45LIBERO project site, Datasets pageOfficial site · 2023 · checked 10 Oct 2026
- 46LIBERO GitHub READMERepository · Mar 2025 · checked 10 Oct 2026
- 47LIBERO branch X-embodiment (unmerged)Repository · 15 Mar 2025 · checked 10 Oct 2026
- 48PyPI: hf-liberoRepository · 10 Jun 2026 · checked 10 Oct 2026
- 49PyPI: liberoRepository · 3 Nov 2025 · checked 10 Oct 2026
- 50huggingface/LIBERO (GitHub fork of upstream; README 'Assets' section)Repository · Jun 2026 · checked 10 Oct 2026
- 51LeRobot pyproject.toml ('libero' extra requires hf-libero)Repository · 2026 · checked 10 Oct 2026
- 52Hugging Face commit list for yifengzhu-hf/LIBERO-datasets (incl. 2025-05-18 file replacement)Repository · 18 May 2025 · checked 10 Oct 2026
- 53LIBERO datasets on Hugging Face (official mirror linked from README), card and file listingRepository · Mar 2025 · checked 10 Oct 2026
- 54LIBERO pull requests and issuesRepository · Oct 2026 · checked 10 Oct 2026
- 55LIBERO requirements.txtRepository · 10 Dec 2024 · checked 10 Oct 2026
- 56LIBERO pull request #146: Support robosuite 1.5.2 / MuJoCo 3.x (open)Repository · Aug 2026 · checked 10 Oct 2026
- 57LIBERO issue #99: compatibility with Python >= 3.9 and new GPUsRepository · Aug 2025 · checked 10 Oct 2026
- 58PyPI: robosuite (release history)Repository · 2026 · checked 10 Oct 2026
- 59LIBERO problem classes (MountedPanda / OnTheGroundPanda wrappers)Repository · 2023 · checked 10 Oct 2026
- 60physical-intelligence/libero dataset card (1,693 episodes, 40 tasks)Repository · Feb 2025 · checked 10 Oct 2026
- 61LeRobot EvalConfig defaults (n_episodes)Repository · 2026 · checked 10 Oct 2026
- 62LIBERO fixed initial-state file (example: libero_spatial, .pruned_init)Repository · 2023 · checked 10 Oct 2026
- 63LIBERO benchmark loader (get_task_init_states reads <task>.pruned_init)Repository · 2023 · checked 10 Oct 2026
- 64LeRobot commit: report success counts and 95% Wilson intervals (#4628)Repository · 16 Sep 2026 · checked 10 Oct 2026
- 65Former Papers with Code LIBERO leaderboard URL (redirects to Hugging Face Trending Papers)Secondary · 10 Oct 2026 · checked 10 Oct 2026
- 66openvla/modified_libero_rlds dataset cardRepository · Sep 2024 · checked 10 Oct 2026
- 67lerobot/libero dataset cardRepository · Jan 2026 · checked 10 Oct 2026
- 68HuggingFaceVLA/libero dataset cardRepository · Sep 2025 · checked 10 Oct 2026
- 69LIBERO download_utils.py (UT Box URLs, Hugging Face repo id)Repository · Mar 2025 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: LIBERO paper (arXiv 2306.03310v2, full text): no real-robot experiments. Project site, docs and README: none. OpenVLA (2406.09246) and OpenVLA-OFT (2502.19645): real-robot and LIBERO results reported separately. RoboArena (2506.18123): cites LIBERO only. VLA-REPLICA (2605.20774): asserts LIBERO-type benchmarks overestimate real performance, no measurement. 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366): related work only. 2026 audit (2606.04233): calls a sim-vs-real ranking test impractical and does not run one. References in LIBERO-PRO, LIBERO-Plus, LIBERO-Para, LIBERO-Safety, LIBERO-RECOVER and RoboVerse also followed. The drafting pass ran web searches (queries not logged). A fresh web search could not be run on 2026-10-10 (shared search budget used up); re-run before publication. Only PolaRiS (2512.16881) reports a paired measurement.
license_assets: Repo LICENSE, README licence table, full repo git tree (1,367 entries; no other licence files), paper NeurIPS checklist items 4(a) and 4(b), paper full text for asset creators (none named), project site, docs, Hugging Face mirror card, lerobot/libero-assets (no card).
objects: Paper text and appendix, project main and Datasets pages, docs Overview and Datasets pages, README.
leaderboard: Project site (main, Datasets, Research pages), GitHub README, papers with code URL /sota/robot-manipulation-on-libero (redirects to Hugging Face Trending Papers), web search for a LIBERO leaderboard.
version: GitHub tags and releases lists (both empty), setup.py, docs header.
top_score (X-VLA trials): X-VLA paper full text: no count of trials or episodes for LIBERO.
top_score (anything above 99.3% four-suite average): Web searches for 2026 LIBERO averages of 99.4% and above; audit tracker per-suite bests. None found; HarnessPAI's 98.1% averages only three suites. A fresh web search on 2026-10-10 could not run (search budget used up); nothing after the audit's 2026-05-21 snapshot was re-searched beyond papers already cited.
demonstrations (original control or recording frequency): LIBERO paper full text and README: no Hz or frequency stated. Derived copies use 10 fps.
issues.i3 (LIBERO-PRO per-perturbation table): LIBERO-PRO v2 HTML: the results table is an image; text gives only qualitative findings and the π0.5 0.38 value. Used SimVLA Table 3 instead.
license_assets (asset origins): HOPE repo tree and README: no per-object name list to match LIBERO's 14 folders.
Change history
- Created from primary sources (full depth). Re-verified the VLGE prior claims for LIBERO.
- Integrated three verifier reports and re-checked every non-confirmed item at the primary source. Fixed: latest_update (data file replaced 2025-05-18), access (UT Box links dead; no checkpoints), LeRobot trial default (docs 10, code 50), robot wrappers, licence sources, used_by wording (lower bound), i1 per-suite bests, Xiaomi wording, openpi protocol notes marked inferred, sim_to_real level to inferred with caveats. Softened i3, i4, i5 and readings r1, r3, r4; added r5. Added i10 (scene identifies task), i11 (asset licences), derived_benchmarks, dataset_downloads, initial-state count, hf-libero fork, Wilson intervals. Kept against verifiers: LIBERO-100 'backgrounds' (paper Figure 1 says so); UW and Princeton as real sites (PolaRiS Figure 6 legend names them); SimVLA size now shows both 0.5B (its abstract) and 0.8B (CORAL). Rejected a verifier case for i7: 96.85 vs 96.9 and 94.15 vs 94.2 are roundings of the same openpi runs. Not re-done: open web search for newer top scores and other sim-to-real studies (search budget used up).
- Published to the Atlas with short display texts. The full texts are unchanged.