VLABench
Simulated single-arm benchmark of 100 language-driven manipulation task types testing common sense, semantics and long-horizon planning.1
Last checked 10 Oct 2026Basic entry14 of 20 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm2
Licence
Unclear3
Comparisons with real robots
No comparison found Unknown
Details
Paper has no real-robot experiments. README 'Preview' note promises real-device deployment with a future release. A June 2026 sim-and-real correlation study (arXiv 2606.10366) cites VLABench but measures VLA-Arena, SIMPLER and REALM, not VLABench.
Details
About
- What it is
- Benchmark Inferred4
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- All 11 authors list School of Computer Science, Fudan University; corresponding authors include Xipeng Qiu2
More
Code is hosted under the OpenMOSS GitHub organisation.
- Released
- 2024-12 (arXiv v1 2024-12-24; README: preview version released 2024/12/25)4
- Version
- Preview; no releases or tags5
More
Tags/releases checked via GitHub API: none.
- Status
- Maintained Inferred8
More
No code commits since 2025-11; datasets updated 2026-07.
Setup
Scoring and access
- Scored by
- Progress score, Success rate, Composite index2
- Leaderboard
- None. Scores are only in papers.9
- Code licence
- MIT (LICENSE.txt); copyright line reads 'Copyright (c) 2024 simpler-env', apparently copied from another project3
- Data licence
- MIT tag on 7 of 10 VLABench Hugging Face datasets incl. VLABench/assets; Apache-2.0 on vlabench_composite_ft_lerobot_video; no tag on vlm_evaluation_v1.0 and vlabench_primitive_ft_lerobot7
- Commercial use
- allowed for code and MIT-tagged data; unclear for untagged datasets and third-party 3D assets Inferred7
More
Paper says some assets were generated with AI tools; not legal advice.
- Published at
- ICCV 2025, pp. 11142-11152 (README: accepted 2025/6/26)1
Sources 9
- 1ICCV 2025 Open Access RepositoryOfficial site · checked 10 Oct 2026
- 2VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks (full text)Official site · Dec 2024 · checked 10 Oct 2026
- 3OpenMOSS/VLABench on GitHub (license)Official site · checked 10 Oct 2026
- 4VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksOfficial site · Dec 2024 · checked 10 Oct 2026
- 5VLABenchOfficial site · checked 10 Oct 2026
- 6OpenMOSS/VLABench on GitHub (commits?per_page=3)Repository · checked 10 Oct 2026
- 7api/datasets on Hugging Face (model)Official site · checked 10 Oct 2026
- 8OpenMOSS/VLABench on GitHub (repository)Official site · checked 10 Oct 2026
- 9OpenMOSS/VLABench on GitHub (file README.md)Repository · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).