VLABench

How to read this picture

Simulated single-arm benchmark of 100 language-driven manipulation task types testing common sense, semantics and long-horizon planning.1

Sources
Last checked 10 Oct 2026Basic entry14 of 20 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm2
Licence
Unclear3

Comparisons with real robots

No comparison found Unknown

Details

Paper has no real-robot experiments. README 'Preview' note promises real-device deployment with a future release. A June 2026 sim-and-real correlation study (arXiv 2606.10366) cites VLABench but measures VLA-Arena, SIMPLER and REALM, not VLABench.

Details

About

What it is
Benchmark Inferred4
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
All 11 authors list School of Computer Science, Fudan University; corresponding authors include Xipeng Qiu2
More

Code is hosted under the OpenMOSS GitHub organisation.

Released
2024-12 (arXiv v1 2024-12-24; README: preview version released 2024/12/25)4
Version
Preview; no releases or tags5
More

Tags/releases checked via GitHub API: none.

Last update
Unknown67
More

README 2025/11/10: new fine-tuned baseline checkpoints; last commit 2025-11-116

Hugging Face dataset VLABench/vlabench_primitive_pretrain_lerobot last modified 2026-07-077

Status
Maintained Inferred8
More

No code commits since 2025-11; datasets updated 2026-07.

Setup

Runs in
Simulation2
Robot
One arm2
Setting
Tabletop, Kitchen, Whole home, Retail or logistics, Office or lab2
More

Tasks are tabletop manipulation set inside these scenes.

Scoring and access

Scored by
Progress score, Success rate, Composite index2
Leaderboard
None. Scores are only in papers.9
Code licence
MIT (LICENSE.txt); copyright line reads 'Copyright (c) 2024 simpler-env', apparently copied from another project3
Data licence
MIT tag on 7 of 10 VLABench Hugging Face datasets incl. VLABench/assets; Apache-2.0 on vlabench_composite_ft_lerobot_video; no tag on vlm_evaluation_v1.0 and vlabench_primitive_ft_lerobot7
Commercial use
allowed for code and MIT-tagged data; unclear for untagged datasets and third-party 3D assets Inferred7
More

Paper says some assets were generated with AI tools; not legal advice.

Published at
ICCV 2025, pp. 11142-11152 (README: accepted 2025/6/26)1

Sources 9

  1. 1ICCV 2025 Open Access RepositoryOfficial site · checked 10 Oct 2026
  2. 2VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks (full text)Official site · Dec 2024 · checked 10 Oct 2026
  3. 3OpenMOSS/VLABench on GitHub (license)Official site · checked 10 Oct 2026
  4. 4VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksOfficial site · Dec 2024 · checked 10 Oct 2026
  5. 5VLABenchOfficial site · checked 10 Oct 2026
  6. 6OpenMOSS/VLABench on GitHub (commits?per_page=3)Repository · checked 10 Oct 2026
  7. 7api/datasets on Hugging Face (model)Official site · checked 10 Oct 2026
  8. 8OpenMOSS/VLABench on GitHub (repository)Official site · checked 10 Oct 2026
  9. 9OpenMOSS/VLABench on GitHub (file README.md)Repository · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).