LIBERO

LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

How to read this picture

130 simulated tasks for one robot arm. Many papers use it to report results for vision-language-action (VLA) models.123

Sources
Last checked 10 Oct 2026Full entry75 of 91 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Checked
Compared once with real robots. Correlation r = 0.66 to 0.70.45
Skill
Handling objects
Robot
One arm2
Franka Emika Panda (simulated)
Used by
551 papers3
1,885 citations
Licence
MIT6
Data licence: sources conflict

What a score here does not tell you Inferred

  1. How well a policy will do on a real robot.Only one study has compared LIBERO scores with real-robot results. It found a correlation of r = 0.66 to 0.70 over five policies, and its authors call this poor.
  2. How well a policy handles new objects or new layouts.The test tasks are the same tasks the policy was trained on. Only the start positions change.
  3. Whether a policy understands the instruction.Studies show that policies can identify the task from the scene or from a task number, without reading the instruction.
ChartPublished scores over time
95% AND ABOVE70809010020252026OpenVLA (7B) · 76.5% · 2024-06π0-FAST · 85.5% · 2025-02π0 · 94.2% · 2025-02OpenVLA-OFT (7B) · 97.1% · 2025-02π0.5 · 96.8% · 2025-09X-VLA (0.9B) · 98% · 2025-10π0.5 + MoH (3B) · 99% · 2025-11Xiaomi-Robotics-0 · 98.7% · 2026-02SimVLA · 98.6% · 2026-02CORAL on SimVLA · 99.3% · 2026-03SimpleVLA-RL on OpenVLA-OFT · 99% · 2025-09πRL on π0.5 (Flow-Noise) · 98.3% · 2025-10OpenVLA 76.5%CORAL on SimVLA 99.3%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.789+10

Comparisons with real robots

StudyResultWhat was comparedDone by
PolaRiS
Dec 2025
Correlation r = 0.66 to 0.704The same 5 policies were scored on LIBERO-90 and on real robots. The real-robot tasks were different from the LIBERO tasks. The study’s authors described the result as “poor”.An independent group

Benchmarks built on LIBERO: LIBERO-PRO, LIBERO-Plus, LIBERO-X, LangGap, LIBERO-Para, Libero-V, LIBERO-Safety, LIBERO-RECOVER.202122+5

Our assessment Opinion

A high score shows that a policy learned these simulated tasks. It says little about how the policy will do on real robots.

Reasoning

A high LIBERO score shows that a policy can learn these simulated tasks from their demonstrations. It does not show general manipulation skill. It is weak evidence about real-robot performance, because the only paired study found a correlation of r = 0.66 to 0.70 over five policies.

Confidence: medium

Do not rank models by differences of one or two points at the top.

Reasoning

Do not rank models by differences of one or two points. At the top of the leaderboard, such differences are often smaller than the variation between random seeds and the differences between the test setups that papers use.

Confidence: high

Read the scores on the harder variants together with the standard score.

Reasoning

The harder variants, such as LIBERO-PRO and LIBERO-Plus, show failures that the standard score does not show. Read their scores together with the standard score. None of the variants has been compared with real-robot results, and models can be tuned to them as well.

Confidence: medium

Most papers now report success on 40 known tasks. Check which protocol a number comes from.

Reasoning

LIBERO was designed to measure how a policy carries knowledge across a stream of new tasks (lifelong learning). In the papers we checked, the headline numbers measure success on 40 known tasks instead. Check which protocol a number comes from.

Confidence: high

LIBERO is still useful for quick, low-cost testing during development.

Reasoning

LIBERO is still useful for fast iteration, debugging and comparison with your own earlier runs. It is open, cheap to run and widely reproduced. The authors of the 2026 audit say the same about the benchmarks they audited.

Confidence: medium

Known problems 11, 1 disputed

  1. Top scores are close to 100%

    Since late 2025, the top average scores have been 98% to 99%. The best scores on single suites reach 99.5% to 100%.131415+6

    Details

    At least five papers since October 2025 report four-suite averages of 98.1 to 99.3%. Per-suite bests are 99.5 to 100% on Spatial, Object and Goal, and 98.8% on Long. Several later papers call LIBERO saturated.

  2. Papers test in different ways

    Papers run 10, 20 or 50 trials per task, train on different data and average over different sets of suites.72930+7

    Details

    Papers differ in trials per task (10, 20 or 50), training data (original or the filtered 1,693 episodes), one policy per suite vs one for all suites vs one expert per task, averages over three or four suites, and RL fine-tuning inside the test simulator. There is no validation split, so picking the best checkpoint on the test suite is allowed; for its own probe, the audit found this adds 0.6 to 4.4 points. Smaller settings matter too: the openpi script notes the environment seed moves objects even with fixed initial states; LeRobot says soft and hard resets give slightly different results and asks authors to pin the dataset revision; SimVLA finds single training choices can outweigh architecture changes.

  3. A small model given only a task number scores close to the best

    A model with 0.09 billion parameters was given only a task number and no instruction. It scored within about one point of the best models on three suites.3

    Details

    A 0.09B probe (DINOv2 encoder + MLP) gets a task ID instead of the instruction. It scored 99.0 / 100.0 / 98.8 / 92.4 on Spatial / Object / Goal / Long, within about one point of the best published result on three suites. It was trained and tuned per suite, and these are the best of several checkpoints scored on the test suite. With no checkpoint selection, its mean was 95.1%. The audit says this shows a high LIBERO score is not on its own evidence of general skill. It does not claim high-scoring policies lack skill.

  4. Scores drop sharply after small camera or start-pose changes

    In LIBERO-Plus, success falls from about 95% to below 30% after small changes to the camera or to the robot's start pose.2133

    Details

    LIBERO-Plus (CVPR 2026) perturbs seven factors. In its single-factor analysis, success drops from about 95% to below 30% under modest camera-viewpoint or robot-initial-state changes. Its released 10,030-task test set dropped tasks that all or most baseline models solved, so its scores are not on the same footing as standard LIBERO scores.

  5. Most claimed improvements are not shown to be statistically significant

    Only 19.8% of 789 claimed improvements on LIBERO can be shown to be statistically significant from the published numbers.3

    Details

    Of 789 previous-best-to-new comparisons on LIBERO (Spatial, Object and Goal pooled), 19.8% are provably significant at the 5% level from public scores. The rest show no improvement, are provably not significant, or cannot be tested because only averages are published. Each comparison uses the previous best as stated in the new paper.

  6. Results change on different computers

    In a test of 10 tasks, changing only the computer's CPU changed the simulation results in all 10.3

    Details

    With policy, seed and initial state fixed, changing only the CPU made the simulator state diverge in 10 of 10 LIBERO tasks for OpenVLA-OFT; in 5 of 10 the difference reached images and actions. Changing only the GPU also caused divergence. The audit argues bitwise determinism should not be the only reproducibility test, and finds aggregate moves small.

  7. The scene often shows which task to do

    Each layout is used for only one task, so a policy can tell the task from the scene. Sources disagree about whether this is true for the Goal suite.32324+1

    Details

    LIBERO draws instructions from a fixed set, one per task, so a task ID can replace language (issues.i2). LangGap says LIBERO assigns only one task per layout. LIBERO-Para says all LIBERO-Goal tasks start from one initial state, so there the instruction is the only cue. The two claims conflict for Goal, which the LIBERO paper describes as same objects and layout with different goals.

  8. Models rely on the exact wording of the instruction

    Without the instruction, success on LIBERO-Goal falls to 0% to 10%. When the instruction is reworded, success falls by 22 to 52 points.212024

    Details

    LIBERO-Plus's text says models largely ignore instructions. Its own blank-instruction test shows success falls sharply for most models when the instruction is removed, to about 0 to 10% on LIBERO-Goal for all six models; only OpenVLA-OFT on Object is unchanged. When the target object in the instruction is swapped, success drops to near zero. LIBERO-PRO finds near-identical trajectories under nonsense instructions. LIBERO-Para (EMNLP 2026) finds paraphrases cut success by 22 to 52 points across seven VLA set-ups, mostly through object synonyms; its authors read this as surface-level matching that disrupts task identification.

  9. High-scoring models fail when object positions or tasks changeDisputed

    In LIBERO-PRO, models that score above 90% on LIBERO drop to 0% when object positions or tasks change. The LIBERO-PRO authors say the models memorised the tasks.2016

    Other view: The 2026 audit agrees that scores drop, but it describes the cause as weak generalisation. When the audit resampled start positions within the training range, scores changed by less than one point.3

    Details

    LIBERO-PRO perturbs objects, positions, instructions and tasks. Its abstract says models above 90% fall to 0.0% in its generalised setting. The drop depends on the perturbation: position and task changes push success to near zero, while object and instruction-meaning changes barely lower it. OpenVLA and π0 fail once an object moves more than 0.2 units; π0.5 holds to about 0.4 units and keeps 0.38 on LIBERO-Goal under position change. Nonsense instructions leave trajectories nearly unchanged. The authors read this as rote memorisation.

    The 2026 audit accepts the drops but disputes the label. It says LIBERO-PRO and LIBERO-Plus change inputs outside the training distribution, so a drop may show weak generalisation rather than overfitting. When it redrew initial states inside the distribution (10,000 rollouts per policy), success moved by under one point (main text: Spatial Forcing −0.62, SimVLA +0.30, LeRobot π0.5 −0.14; its appendix table prints the opposite signs). It notes rollout noise leaves each sign uncertain.

  10. The same model gets different scores in different papers

    Different papers report the score of the same model, π0, as 86.0, 94.15 and 96.8.81030+8

    Details

    π0: 94.15 (openpi, 50 trials per task by script default), 86.0 (SmolVLA paper, 10 trials per task), 96.8 (T-MEE paper's own run). π0.5: 96.85 (openpi), 97.5 (LeRobot, 10 episodes per task), 97.7 (MoH paper's run). OpenVLA-OFT: 97.1 (its paper) vs 91.0 (SimpleVLA-RL's re-implementation with one camera and no proprioception; πRL reprints it labelled only 'OpenVLA-OFT'). SimVLA's own text and Table 2 disagree on two suites.

  11. The included 3D models have no stated licence

    The 3D models that come with LIBERO have no licence file. Some of them appear to come from collections with restricted licences. Inferred343536+2

    Details

    The repo and Hugging Face's re-hosted copy ship 3D assets with no licence file. Folder names suggest third-party origins. stable_hope_objects holds 14 grocery items; NVIDIA's HOPE set is 28 toy grocery items with 3D meshes under CC BY-NC-SA 4.0. turbosquid_objects holds 17 models; TurboSquid's Royalty Free License bars giving away model files outside a permitted Creation. stable_scanned_objects (11 items) has unchecked origin.

Details

About

What it is
Benchmark1
More

The paper calls it a benchmark with fixed task suites, demonstrations and metrics.

Built by
The University of Texas at Austin, Sony AI, Tsinghua University13839
More

Authors: Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, Peter Stone. The project site's Research page omits Yihao Feng; arXiv, NeurIPS and the README list all seven.

The University of Texas at Austin · LARG, RPL, and Statistical Learning & AI groups (project site footer). Six of seven authors.138

Sony AI · Second affiliation of Peter Stone.1

Tsinghua University · Affiliation of Chongkai Gao.1

Released
June 2023, at NeurIPS 20234041
More

arXiv v1 on 2023-06-05. Published at NeurIPS 2023, Datasets and Benchmarks Track.

The GitHub repository was created on 2023-04-08, before the paper. arXiv v2 is dated 2023-10-14.

Version
0.1.0. There are no tagged releases.424344
More

Upstream package version 0.1.0. No tagged releases. Four suites: LIBERO-Spatial, LIBERO-Object, LIBERO-Goal (10 tasks each) and LIBERO-100 (split into LIBERO-90 and LIBERO-10).

setup.py and the docs both say 0.1.0. The GitHub tags and releases lists are empty. Two other upstream branches exist: 'docs' (last commit 2023-10-13) and 'X-embodiment' (see items). The packages on PyPI are Hugging Face builds, not upstream (see items).

LIBERO-Spatial · 10 tasks. Same objects; two identical bowls in different places. Varies spatial layout. Task: put the named bowl on the plate.17

LIBERO-Object · 10 tasks. Same layout; a different object to pick and place in each task. Varies object type.17

LIBERO-Goal · 10 tasks. Same objects, same layout. Varies the goal (the motion or behaviour asked for).17

LIBERO-100 · 100 tasks with mixed objects, layouts and backgrounds. Split into LIBERO-90 (90 short tasks, used for pretraining) and LIBERO-10 (10 long tasks, used for evaluation).145

LIBERO-10 = LIBERO-Long · Two names for the same 10 long-horizon tasks. The paper says LIBERO-Long; the code says libero_10.1467

Branch X-embodiment (unmerged) · Cross-embodiment code (first commit 2024-08-22) and a robosuite 1.5 update (2025-03-15). Never merged into master.47

hf-libero 0.1.4 (Hugging Face fork) · PyPI hf-libero, latest 0.1.4 on 2026-06-10, built from huggingface/LIBERO, a GitHub fork of upstream. LeRobot's 'libero' extra requires it. The fork downloads 3D assets at run time from the Hugging Face dataset lerobot/libero-assets. PyPI 'libero' 0.1.1 (2025-11-03) points to huggingface/lerobot-libero, now archived.484950+2

Last update
May 2025. One data file was replaced.525344
More

2025-05-18: one LIBERO-90 demonstration file replaced on the official Hugging Face mirror, with no changelog. Last code commit: 2025-03-15 (Hugging Face download support).

Replaced file: libero_90/LIVING_ROOM_SCENE1_pick_up_the_tomato_sauce_and_put_it_in_the_basket_demo.hdf5. Commit message: the original file 'seems to be corrupted'. Size 590,254,080 bytes before, 806,284,784 after (Hub tree at both revisions). Copies made before that date may hold the old file. The mirror was created 2025-03-13; its card was last edited 2025-03-17.

Status
No code changes since March 2025. Still widely used. Inferred445452+5
More

Upstream code unchanged since 2025-03-15; last data fix 2025-05-18. A Hugging Face fork is maintained. Use is very active.

Last merged pull request 2025-01-03. 108 open items (24 PRs, 84 issues); community PRs from October 2026 are unmerged. Last comment by a repo collaborator: 2025-08-06 (issue #99, Python >= 3.9 and new GPUs). requirements.txt pins robosuite 1.4.0 (released 2022-12-01); robosuite on PyPI is 1.5.2. PR #146 for robosuite 1.5.2 / MuJoCo 3.x is open. The hf-libero fork (0.1.4, 2026-06-10) keeps LIBERO installable for LeRobot users. For current use, see facts.used_by.

Setup

Runs in
Simulation1
Simulator
robosuite 1.4, built on MuJoCo1552
More

Paper: built on robosuite. requirements.txt pins robosuite==1.4.0. Environment code renders with MuJoCo. Tasks are written as BDDL/PDDL files.

Robot
One arm2
Robot model
Franka Emika Panda (simulated)25911
More

Code default is robots=['Panda']. Tabletop and kitchen problem classes wrap it as MountedPanda; the floor problem (LIBERO-Object) wraps it as OnTheGroundPanda. The paper text does not name the robot; OpenVLA-OFT names it.

Setting
Tabletop, Kitchen, Whole home Inferred134
More

Living room and study are mapped to 'home'. These are single workspaces, not whole houses.

Tasks
130 tasks in 4 suites40134
More

130 tasks in four suites (10 + 10 + 10 + 100)

Repo has 10 + 10 + 10 + 90 + 10 BDDL task files (counted 2026-10-10). Most VLA papers use only the 40 tasks of Spatial, Object, Goal and Long.

Scenes
20 layouts in LIBERO-100 Inferred153
More

20 scene layouts in LIBERO-100: 10 kitchen, 6 living room, 4 study

Counted by us from the scene names in the paper's LIBERO-90 task table (Appendix E.3) and the dataset file names. LIBERO-10 reuses scenes from that set. Spatial, Object and Goal use their own layouts and are not counted here.

Training data
50 human demonstrations per task Inferred145
More

50 human demonstrations per task; 6,500 in total by our arithmetic. The project site says 65,000.

CONFLICT. The paper states 50 trajectories per task, collected by human teleoperation with a 3Dconnexion SpaceMouse. 130 x 50 = 6,500. The project Datasets page states 65,000. OpenVLA also counts 500 demonstrations per 10-task suite, which supports 50 per task.

50 per task · Human teleoperation with a 3Dconnexion SpaceMouse1

Observation content · Workspace and wrist RGB cameras, proprioception, language instruction, PDDL scene description45

128 x 128 px images · Original image resolution. OpenVLA re-rendered all demonstrations at 256 x 256.7

1,693 episodes (filtered) · OpenVLA replayed demonstrations and dropped failures: 68, 46, 72 and 121 of 500 in Spatial, Object, Goal and Long. The widely used Physical Intelligence copy has 1,693 episodes, 273,465 frames at 10 fps, over 40 tasks.76019

about 100.4 GB · Total size of the 130 HDF5 files on the official Hugging Face mirror (decimal GB, 10^9 bytes)53

Changes at test
Only the start positions change Inferred46720
More

In the common VLA protocol, test tasks are the training tasks. Only the initial object placement differs, drawn from a fixed set per task.

README: initial states are fixed for benchmarking. OpenVLA keeps 'the same initial environment configurations' as the original benchmark. In the original lifelong protocol, suites isolate shifts in layout, object and goal between tasks, but each new task comes with its own demonstrations, so this is transfer, not zero-shot generalisation.

Scoring and access

Scored by
Success rate17
More

Success is a binary goal predicate checked by the simulator. The original lifelong metrics (FWT, NBT, AUC) are built from success rates and have no taxonomy value of their own.

Score
Success rate for each suite, averaged across suites1712
More

Original paper: three lifelong-learning metrics, forward transfer (FWT), negative backward transfer (NBT) and area under the success curve (AUC), computed while tasks arrive one by one. Later VLA papers: train on the demonstrations, then report the success rate per suite and the plain average over Spatial, Object, Goal and Long.

The two protocols answer different questions. The first asks how well a learner keeps and transfers skills over a sequence. The second asks how well one model fits 40 known tasks. Almost all headline numbers today use the second.

Trials
10 to 50 per task, depending on the paper1710+5
More

Fixed initial states: 50 per task · The benchmark loads a fixed set of initial states per task. This caps the common protocol at 50 distinct starts per task, 500 per suite.636219+1

Original paper · 20 rollouts per task, max 600 steps, every 5 epochs; 3 seeds (100, 200, 300)1

OpenVLA convention · 500 trials per suite (50 per task) from the fixed initial states; averaged over 3 seeds. Third-person images rotated 180 degrees at train and test time, because the environments rendered upside down on their hardware.7

openpi evaluation script · 50 trials per task, seed 7, images resized to 224 px, step limits 220 to 520 by suite. A code comment warns the environment seed changes object positions even with a fixed initial state.10

LeRobot · Docs recommend 10 episodes per task (400 in total) and averaging over 3 seeds. The code default is 50 episodes (EvalConfig.n_episodes).2961

SmolVLA paper · 10 trials per task30

NVIDIA GR00T N1.7 example · 200 episodes per suite (20 per task)31

Who runs it
Each team tests its own model Inferred3846
More

No organiser runs submissions. Each paper runs its own evaluation.

Error bars
Sometimes reported Inferred173+2
More

The original paper reports mean and standard error over 3 seeds. OpenVLA reports standard error. The 2026 audit finds most benchmarks report only an aggregate success rate. Since 2026-09-16, LeRobot's evaluator prints a 95% Wilson interval and success counts next to every success rate.

Leaderboard
None. Scores are only in papers. Inferred384665
More

No leaderboard on the project site or in the repo. The former Papers with Code LIBERO page now redirects to Hugging Face Trending Papers (HTTP 302, checked 2026-10-10).

Code licence
MIT6
More

LICENSE file: MIT, copyright 2023 Lifelong Robot Learning.

Data licence
Sources conflict: CC BY 4.0 or Apache-2.04653
More

Conflict: the README says CC BY 4.0; the official Hugging Face dataset card says Apache-2.0. The Hugging Face copy is now the only live official download.

CONFLICT between two primary sources from the same team. The README links the Hugging Face mirror as an official download, and download_utils.py uses it. Derived copies by other groups carry their own labels (items); they do not decide LIBERO's own licence.

MIT · openvla/modified_libero_rlds (OpenVLA's filtered copy)66

CC-BY-4.0 · physical-intelligence/libero (1,693 episodes)60

Apache-2.0 · lerobot/libero (recommended by LeRobot docs)67

Apache-2.0 · HuggingFaceVLA/libero68

Asset licence
Unknown
More

The repo ships 3D assets in folders named turbosquid_objects, stable_hope_objects and stable_scanned_objects. The full repo tree has no licence file besides the root MIT LICENSE, which covers 'the Software'. In the paper's NeurIPS checklist, 4(a) 'did you cite the creators' is Yes, 4(b) on asset licences is N/A; the paper text names no asset source. Hugging Face's re-hosted asset copy (lerobot/libero-assets) has no card and no licence. See issues.i11 for likely origins.

Access
Open. The data is on Hugging Face.534645+1
More

Code on GitHub. Data only from the ungated Hugging Face dataset yifengzhu-hf/LIBERO-datasets. The four original UT Box links (project Datasets page and download_utils.py) return HTTP 404 (checked 2026-10-10). The Datasets page's four 'Best Model Checkpoints' links are empty, so no official checkpoints can be downloaded. No registration needed.

Commercial use
Unclear Inferred64653+2
More

Code (MIT) and data (CC BY 4.0 or Apache-2.0, whichever applies) both allow commercial use with attribution. The bundled third-party 3D assets have no stated licence, and some folder names point to sources with non-commercial or no-redistribution terms (issues.i11). So the whole package is unclear. Not legal advice.

Sources 69

  1. 1LIBERO paper, full text v2Paper · Oct 2023 · checked 10 Oct 2026
  2. 2LIBERO environment wrapper source (robot and renderer)Repository · 2023 · checked 10 Oct 2026
  3. 3What Are We Actually Benchmarking in Robot Manipulation?Paper · Jun 2026 · checked 10 Oct 2026
  4. 4PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Section 5.2, Figures 6 and 7, Appendix C.1)Paper · Dec 2025 · checked 10 Oct 2026
  5. 5VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of VLA ModelsPaper · May 2026 · checked 10 Oct 2026
  6. 6LIBERO LICENSE fileRepository · Jun 2023 · checked 10 Oct 2026
  7. 7OpenVLA: An Open-Source Vision-Language-Action Model (Appendix E, LIBERO)Paper · Jun 2024 · checked 10 Oct 2026
  8. 8openpi LIBERO example README at first commit (π0 and π0-FAST results)Repository · 4 Feb 2025 · checked 10 Oct 2026
  9. 9openpi training configs (pi0_libero, pi05_libero use physical-intelligence/libero)Repository · 2025 · checked 10 Oct 2026
  10. 10openpi LIBERO evaluation scriptRepository · 2025 · checked 10 Oct 2026
  11. 11Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)Paper · Feb 2025 · checked 10 Oct 2026
  12. 12openpi LIBERO example README (current: π0.5 results)Repository · Sep 2025 · checked 10 Oct 2026
  13. 13X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelPaper · Oct 2025 · checked 10 Oct 2026
  14. 14Mixture of Horizons in Action Chunking (MoH)Paper · Nov 2025 · checked 10 Oct 2026
  15. 15Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time ExecutionPaper · Feb 2026 · checked 10 Oct 2026
  16. 16SimVLA: A Simple VLA Baseline for Robotic ManipulationPaper · Feb 2026 · checked 10 Oct 2026
  17. 17CORAL: Scalable Multi-Task Robot Learning via LoRA ExpertsPaper · Mar 2026 · checked 10 Oct 2026
  18. 18SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningPaper · Sep 2025 · checked 10 Oct 2026
  19. 19πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action ModelsPaper · Oct 2025 · checked 10 Oct 2026
  20. 20LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (v2, 2026-05-25)Paper · Oct 2025 · checked 10 Oct 2026
  21. 21LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action ModelsPaper · Oct 2025 · checked 10 Oct 2026
  22. 22LIBERO-X: Robustness Litmus for Vision-Language-Action ModelsPaper · Feb 2026 · checked 10 Oct 2026
  23. 23LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action ModelsPaper · Feb 2026 · checked 10 Oct 2026
  24. 24LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA ModelsPaper · Mar 2026 · checked 10 Oct 2026
  25. 25VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial ModelingPaper · Jun 2026 · checked 10 Oct 2026
  26. 26LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action ModelsPaper · Jun 2026 · checked 10 Oct 2026
  27. 27LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation ModelsPaper · Sep 2026 · checked 10 Oct 2026
  28. 28Reshaping Action Error Distributions for Reliable Vision-Language-Action Models (T-MEE)Paper · Feb 2026 · checked 10 Oct 2026
  29. 29LeRobot documentation: LIBERORepository · 2026 · checked 10 Oct 2026
  30. 30SmolVLA: A vision-language-action model for affordable and efficient roboticsPaper · Jun 2025 · checked 10 Oct 2026
  31. 31Isaac-GR00T LIBERO example and results (GR00T N1.7)Repository · Jul 2026 · checked 10 Oct 2026
  32. 32HarnessPAI: An Evolving Harness for Physical AI (three-suite average)Paper · Sep 2026 · checked 10 Oct 2026
  33. 33LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models (CVPR 2026 poster page)Paper · Jun 2026 · checked 10 Oct 2026
  34. 34LIBERO repo tree: BDDL task files and asset foldersRepository · Mar 2025 · checked 10 Oct 2026
  35. 35lerobot/libero-assets (Hugging Face dataset; no card)Repository · Nov 2025 · checked 10 Oct 2026
  36. 36NVIDIA HOPE dataset README (licence section)Repository · 2021 · checked 10 Oct 2026
  37. 37TurboSquid Royalty Free LicenseOfficial site · unknown · checked 10 Oct 2026
  38. 38LIBERO project site, main pageOfficial site · 2023 · checked 10 Oct 2026
  39. 39LIBERO project site, Research pageOfficial site · 2023 · checked 10 Oct 2026
  40. 40LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning (arXiv abstract page)Paper · Jun 2023 · checked 10 Oct 2026
  41. 41LIBERO, NeurIPS 2023 proceedings pagePaper · Dec 2023 · checked 10 Oct 2026
  42. 42LIBERO setup.pyRepository · 2023 · checked 10 Oct 2026
  43. 43LIBERO documentation: DatasetsOfficial site · 2023 · checked 10 Oct 2026
  44. 44LIBERO commit history, tags and releasesRepository · 15 Mar 2025 · checked 10 Oct 2026
  45. 45LIBERO project site, Datasets pageOfficial site · 2023 · checked 10 Oct 2026
  46. 46LIBERO GitHub READMERepository · Mar 2025 · checked 10 Oct 2026
  47. 47LIBERO branch X-embodiment (unmerged)Repository · 15 Mar 2025 · checked 10 Oct 2026
  48. 48PyPI: hf-liberoRepository · 10 Jun 2026 · checked 10 Oct 2026
  49. 49PyPI: liberoRepository · 3 Nov 2025 · checked 10 Oct 2026
  50. 50huggingface/LIBERO (GitHub fork of upstream; README 'Assets' section)Repository · Jun 2026 · checked 10 Oct 2026
  51. 51LeRobot pyproject.toml ('libero' extra requires hf-libero)Repository · 2026 · checked 10 Oct 2026
  52. 52Hugging Face commit list for yifengzhu-hf/LIBERO-datasets (incl. 2025-05-18 file replacement)Repository · 18 May 2025 · checked 10 Oct 2026
  53. 53LIBERO datasets on Hugging Face (official mirror linked from README), card and file listingRepository · Mar 2025 · checked 10 Oct 2026
  54. 54LIBERO pull requests and issuesRepository · Oct 2026 · checked 10 Oct 2026
  55. 55LIBERO requirements.txtRepository · 10 Dec 2024 · checked 10 Oct 2026
  56. 56LIBERO pull request #146: Support robosuite 1.5.2 / MuJoCo 3.x (open)Repository · Aug 2026 · checked 10 Oct 2026
  57. 57LIBERO issue #99: compatibility with Python >= 3.9 and new GPUsRepository · Aug 2025 · checked 10 Oct 2026
  58. 58PyPI: robosuite (release history)Repository · 2026 · checked 10 Oct 2026
  59. 59LIBERO problem classes (MountedPanda / OnTheGroundPanda wrappers)Repository · 2023 · checked 10 Oct 2026
  60. 60physical-intelligence/libero dataset card (1,693 episodes, 40 tasks)Repository · Feb 2025 · checked 10 Oct 2026
  61. 61LeRobot EvalConfig defaults (n_episodes)Repository · 2026 · checked 10 Oct 2026
  62. 62LIBERO fixed initial-state file (example: libero_spatial, .pruned_init)Repository · 2023 · checked 10 Oct 2026
  63. 63LIBERO benchmark loader (get_task_init_states reads <task>.pruned_init)Repository · 2023 · checked 10 Oct 2026
  64. 64LeRobot commit: report success counts and 95% Wilson intervals (#4628)Repository · 16 Sep 2026 · checked 10 Oct 2026
  65. 65Former Papers with Code LIBERO leaderboard URL (redirects to Hugging Face Trending Papers)Secondary · 10 Oct 2026 · checked 10 Oct 2026
  66. 66openvla/modified_libero_rlds dataset cardRepository · Sep 2024 · checked 10 Oct 2026
  67. 67lerobot/libero dataset cardRepository · Jan 2026 · checked 10 Oct 2026
  68. 68HuggingFaceVLA/libero dataset cardRepository · Sep 2025 · checked 10 Oct 2026
  69. 69LIBERO download_utils.py (UT Box URLs, Hugging Face repo id)Repository · Mar 2025 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: LIBERO paper (arXiv 2306.03310v2, full text): no real-robot experiments. Project site, docs and README: none. OpenVLA (2406.09246) and OpenVLA-OFT (2502.19645): real-robot and LIBERO results reported separately. RoboArena (2506.18123): cites LIBERO only. VLA-REPLICA (2605.20774): asserts LIBERO-type benchmarks overestimate real performance, no measurement. 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366): related work only. 2026 audit (2606.04233): calls a sim-vs-real ranking test impractical and does not run one. References in LIBERO-PRO, LIBERO-Plus, LIBERO-Para, LIBERO-Safety, LIBERO-RECOVER and RoboVerse also followed. The drafting pass ran web searches (queries not logged). A fresh web search could not be run on 2026-10-10 (shared search budget used up); re-run before publication. Only PolaRiS (2512.16881) reports a paired measurement.

license_assets: Repo LICENSE, README licence table, full repo git tree (1,367 entries; no other licence files), paper NeurIPS checklist items 4(a) and 4(b), paper full text for asset creators (none named), project site, docs, Hugging Face mirror card, lerobot/libero-assets (no card).

objects: Paper text and appendix, project main and Datasets pages, docs Overview and Datasets pages, README.

leaderboard: Project site (main, Datasets, Research pages), GitHub README, papers with code URL /sota/robot-manipulation-on-libero (redirects to Hugging Face Trending Papers), web search for a LIBERO leaderboard.

version: GitHub tags and releases lists (both empty), setup.py, docs header.

top_score (X-VLA trials): X-VLA paper full text: no count of trials or episodes for LIBERO.

top_score (anything above 99.3% four-suite average): Web searches for 2026 LIBERO averages of 99.4% and above; audit tracker per-suite bests. None found; HarnessPAI's 98.1% averages only three suites. A fresh web search on 2026-10-10 could not run (search budget used up); nothing after the audit's 2026-05-21 snapshot was re-searched beyond papers already cited.

demonstrations (original control or recording frequency): LIBERO paper full text and README: no Hz or frequency stated. Derived copies use 10 fps.

issues.i3 (LIBERO-PRO per-perturbation table): LIBERO-PRO v2 HTML: the results table is an image; text gives only qualitative findings and the π0.5 0.38 value. Used SimVLA Table 3 instead.

license_assets (asset origins): HOPE repo tree and README: no per-object name list to match LIBERO's 14 folders.

Change history

  1. Created from primary sources (full depth). Re-verified the VLGE prior claims for LIBERO.
  2. Integrated three verifier reports and re-checked every non-confirmed item at the primary source. Fixed: latest_update (data file replaced 2025-05-18), access (UT Box links dead; no checkpoints), LeRobot trial default (docs 10, code 50), robot wrappers, licence sources, used_by wording (lower bound), i1 per-suite bests, Xiaomi wording, openpi protocol notes marked inferred, sim_to_real level to inferred with caveats. Softened i3, i4, i5 and readings r1, r3, r4; added r5. Added i10 (scene identifies task), i11 (asset licences), derived_benchmarks, dataset_downloads, initial-state count, hf-libero fork, Wilson intervals. Kept against verifiers: LIBERO-100 'backgrounds' (paper Figure 1 says so); UW and Princeton as real sites (PolaRiS Figure 6 legend names them); SimVLA size now shows both 0.5B (its abstract) and 0.8B (CORAL). Rejected a verifier case for i7: 96.85 vs 96.9 and 94.15 vs 94.2 are roundings of the same openpi runs. Not re-done: open web search for newer top scores and other sim-to-real studies (search budget used up).
  3. Published to the Atlas with short display texts. The full texts are unchanged.