RLBench
RLBench: The Robot Learning Benchmark & Learning Environment
RLBench is a set of 100 simulated tasks for one Franka Panda robot arm. Papers on robot manipulation policies (the robot's control models) use it, and most report results on an 18-task subset.12
What a score here does not tell you Inferred
- How well a policy will do on a real robot.No study has compared RLBench and real-robot scores for the same policies.
- How well a policy handles new objects.The common protocol tests only objects seen in training.
- How well a policy copes with changes to the scene.THE COLOSSEUM, a test built on RLBench, shows large drops in success under scene changes.
Comparisons with real robots
Tried on real robots, not compared Inferred125+4
Details
The RLBench paper offers domain randomisation and swappable arms for sim-to-real research but has no real-robot experiment. The closest evidence is THE COLOSSEUM (RSS 2024), built on RLBench: it mirrored 4 tasks with 3D-printed objects on a real Franka, trained one PerAct on 5 real demonstrations per task and another in simulation, and tested each under matching perturbations (10 episodes × 3 runs). Per-factor R² between the two ranged from 0.46 to 0.94 (mean 0.614). These are two separately trained models of one method compared across perturbation conditions, so we do not count it as a paired evaluation of the same policies. PolaRiS and the 2026 sim-real recipe paper mention RLBench only in related work.
Our assessment Opinion
Check which RLBench subset and protocol a number uses.
Reasoning
When a paper reports 'RLBench', it almost always means PerAct's 18-task protocol, not the 100-task benchmark or its few-shot challenge. Check the task subset, the number of demonstrations and the inputs before comparing.
Confidence: high
A high score shows that a policy fits 18 known tasks. It does not show that the policy copes with scene changes.
Reasoning
A high RLBench-18 score shows a policy fits 18 known tasks with seen objects in one scene. Scene perturbations cut scores sharply, and no study has checked RLBench scores against real-robot results for the same policies.
Confidence: medium
Ignore gaps of a few points between top results.
Reasoning
Gaps of a few points at the top are not reliable. Data-loader details move the same model by several points, and papers re-report earlier models with different numbers.
Confidence: high
Commercial users need permission from Imperial College London. The default edition of the simulator also excludes commercial use.
Reasoning
Companies should read the licence before using RLBench. It and the default simulator edition exclude commercial use and product research without permission. This is our reading, not legal advice.
Confidence: high
Known problems 7
The licence is non-commercial and the simulator edition is for education only
The licence does not allow commercial use. The default simulator edition is for schools and universities only.11203+1
Details
RLBench's custom licence allows only non-commercial, internal or academic research, bans transfer and sub-licensing, lets Imperial terminate at any time, and asks publications to credit Imperial and send it a copy. The README installs CoppeliaSim Edu, whose terms exclude companies, research institutions and commercial use. Main forks keep the custom licence (GitHub shows NOASSERTION).
Maintenance has stopped and 3D assets lack licences
The code has had no functional changes since July 2024 and uses an old simulator version. The third-party 3D models have no licence information.22320+1
Details
The last functional commit is from July 2024 and the code pins CoppeliaSim 4.1.0 (the vendor's current version is 4.10.0). 94 issues are open. The README credits 3D models to five third-party model sites without per-model licence information.
The same model gets different scores
The score of one model, RVT-2, falls from 81.4% to 74.1% after one change to how training data is loaded.479+2
Details
PerAct's paper gives no average; RVT re-evaluated its released model at 49.4%. BridgeVLA reports 88.2%; its follow-up BridgeVLA++ reports BridgeVLA at 90.5 ± 1.1%. ARP found that RVT-2's 81.4% depends on a data-loader detail inherited from C2F-ARM: removing a randomised timestep gives 77.0%, and giving correct timesteps gives 74.1%. RVT notes that earlier baselines picked the best model per task on validation sets, which overstates their multi-task scores.
Results vary from one evaluation run to the next
The motion planner makes random choices, so results vary between runs. Some papers report only one run.42317
Details
RVT runs each model five times on the same 25 episodes per task because the sampling-based motion planner is random. PerAct distributes fixed pre-generated datasets because data generation samples scenes randomly. Some papers report one evaluation run.
Scores fall when the scene changes
In THE COLOSSEUM, success drops by 30 to 50% for each type of scene change, and by at least 75% when changes are combined.17
Details
THE COLOSSEUM perturbs 20 RLBench tasks along 14 factors such as colour, texture, lighting, distractors and camera pose. It reports success dropping 30 to 50% per factor and at least 75% when factors are combined. PerAct trained without perturbations fell from 34.5% to 6.4% with all perturbations on.
Most RLBench results come from a modified 18-task subset
Most papers report results on a modified subset of 18 tasks defined by PerAct, instead of the full 100-task benchmark.2241+1
Details
Most recent papers report PerAct's RLBench-18 protocol: 18 tasks, some modified to add variations (249 in total), 100 demonstrations and 25 test episodes per task, four 128 × 128 RGB-D cameras and keyframe actions. Others use 74 tasks (Hiveformer) or their own subsets. The paper's own few-shot challenge is rarely reported, and the paper (10% of tasks held out) and README (5 test tasks) disagree on its split.
Details
About
- What it is
- Benchmark1
More
The paper presents RLBench as a benchmark and learning environment with 100 tasks and a few-shot challenge.
- Built by
- Imperial College London (Dyson Robotics Lab)1
More
Authors: Stephen James, Zicong Ma, David Rovick Arrojo, Andrew J. Davison (Dyson Robotics Lab and UROP, Imperial College London). The research was supported by Dyson Technology Ltd.
- Released
- September 2019, in IEEE Robotics and Automation Letters (RA-L) 202025263+1
More
arXiv v1 on 2019-09-26 (repository created the same day). Published in IEEE Robotics and Automation Letters 5(2), pp. 3019-3026 (April 2020), presented at ICRA 2020.
- Version
- 1.2.0. The last release tag is 1.1.0.27283
More
1.2.0 in the code; the last GitHub release tag is 1.1.0 (2021-05-31). Not on PyPI; installed from GitHub.
RLBench-18 (PerAct protocol) · 18 tasks, some modified to add variations (249 variations in total), 100 training demos and 25 test episodes per task. Defined by PerAct (CoRL 2022), not by the RLBench authors. Used by most recent papers.2
RLBench-74 (Hiveformer protocol) · 74 tasks, CoRL 2022.2429
Few-shot task sets · FS10_V1, FS25_V1, FS50_V1, FS95_V1: 10 to 95 training tasks plus 5 test tasks. Multi-task sets MT15_V1 to MT100_V1.330
- Last update
- January 2025. The last commit fixed a typo.223
More
Last commit 2025-01-25 (typo fix in a tutorial). Last functional changes 2024-07-02/03 (Gymnasium support, arm velocity and acceleration limits).
README announcements: v1.2.0 on 2022-02-18 (breaking action-mode API changes); shaped rewards for two tasks on 2022-05-11.
- Status
- No functional changes since July 2024. Still used in new papers. Inferred221231
More
No functional change since July 2024; last commit January 2025. Still used in new papers.
94 open issues. Issue #268 (2025-01, open): the bulb holder is missing from generated light_bulb_in demonstrations; the reporter switched tasks.
Setup
- Runs in
- Simulation1
- Simulator
- CoppeliaSim 4.1, through PyRep3132
More
The README downloads CoppeliaSim Edu 4.1.0 for Ubuntu 20.04. The CoppeliaSim website lists V4.10.0 as current. PyRep is installed from GitHub; the PyPI name 'pyrep' belongs to an unrelated AGPL package.
- Robot
- One arm3
- Robot model
- Franka Emika Panda (simulated)31
More
Mico, Jaco, Sawyer and UR5 can be swapped in, but the README says the arm should remain the Franka Panda for benchmarking.
- Setting
- Tabletop1
- Tasks
- 100 tasks, of which 18 are commonly used130
More
100 tasks in the paper; 106 task files in the repository today
Repository count by us: .py files in rlbench/tasks excluding __init__.py (2026-10-10). Recent papers mostly use the 18-task subset.
- Scenes
- 1 scene1
More
1 shared table scene
- Training data
- Unlimited, generated by a motion planner123
More
Unlimited demonstrations generated by a motion planner (OMPL) from task waypoints; lengths 100 to 1,000 steps. No fixed official dataset.
PerAct pre-generated data · Train 100, validation 25 and test 25 episodes per task for the 18 tasks, about 116 GB on Google Drive. PerAct says using them helps reproducibility because data generation samples scenes randomly.23
- Changes at test
- New object positions. The objects are the same as in training.21
More
In the common RLBench-18 protocol, test episodes use new object poses and sampled variations (colours, sizes, targets) that also appear in training. Train and test objects are the same.
PerAct states that it does not test generalisation to unseen objects. The original few-shot challenge (new tasks) is rarely reported; derived benchmarks such as GemBench and AGNOSTOS test new tasks.
Scoring and access
- Scored by
- Success rate12
More
Each task has a sparse reward of +1 on completion, checked by task-specific success conditions.
- Score
- Success rate over 18 tasks21
More
In the RLBench-18 protocol: one multi-task policy, each test episode scored 0 or 100, averaged over 25 episodes per task and over the 18 tasks. The paper's own challenge asks for 1-, 5- and 20-shot success on held-out tasks.
- Trials
- 25 per task, often repeated in 5 runs249
More
PerAct · 25 episodes per task (450 in total), at most 25 keyframe steps; best checkpoint chosen on a separate 25-episode validation set.2
RVT and later · Each model run 5 times on the same 25 episodes per task, because the sampling-based motion planner is random; mean and standard deviation reported.4
BridgeVLA++ · Mean ± std over five random seeds, 25 evaluation episodes per seed.9
- Who runs it
- Each team tests its own model Inferred33334
More
No submission process or organiser-run evaluation on the site or README.
- Error bars
- Sometimes reported Inferred246+2
More
RVT, RVT-2, 3D Diffuser Actor, SAM2Act, BridgeVLA and BridgeVLA++ report standard deviations over 5 runs. PerAct reports one evaluation; COLOSSEUM reports one training seed and one evaluation seed.
- Leaderboard
- None. Scores are only in papers.33334
More
No leaderboard on the project site or README. In issue #171 (2022) the maintainer said no official baselines exist because simple policies did not work on these sparse-reward tasks.
- Code licence
- Custom licence, non-commercial use only1112
More
Custom licence: free, non-exclusive, non-transferable; use only for non-commercial, internal or academic research. BSD terms for BSD elements.
Key clauses: 2(b) bans commercial use, including research to develop products for sale and paid services; commercial use requires contacting researchcontracts.engineering@imperial.ac.uk. 2(d) bans transfer and sub-licensing. 2(f) requires publications to credit the software as licensed from Imperial and to send Imperial a copy. 7(a) Imperial may terminate at any time; 7(d) restrictions expire 10 years after first use. English law. GitHub shows 'NOASSERTION'.
- Data licence
- not-applicable Inferred123
More
No official dataset is distributed; users generate demonstrations. PerAct's pre-generated data are produced with RLBench and hosted by the PerAct authors.
- Asset licence
- Unknown
More
The README says 3D models came from turbosquid.com, cgtrader.com, free3d.com, thingiverse.com and cadnav.com. No per-model licence or attribution file was found in the README or LICENSE; the RLBench licence covers 'the Software'.
- Access
- Open. Downloading the code binds the user to the licence.31120
More
Code on GitHub with no registration; downloading binds the user to the licence. The simulator download requires accepting CoppeliaSim's terms.
Sources 34
- 1RLBench paper, full text v1Paper · Sep 2019 · checked 10 Oct 2026
- 2Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation (PerAct)Paper · Sep 2022 · checked 10 Oct 2026
- 3RLBench GitHub README (install, task sets, announcements, acknowledgements)Repository · 2024 · checked 10 Oct 2026
- 4RVT: Robotic View Transformer for 3D Object ManipulationPaper · Jun 2023 · checked 10 Oct 2026
- 5RVT-2: Learning Precise Manipulation from Few DemonstrationsPaper · Jun 2024 · checked 10 Oct 2026
- 6SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic ManipulationPaper · Jan 2025 · checked 10 Oct 2026
- 7BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language ModelsPaper · Jun 2025 · checked 10 Oct 2026
- 8ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic ManipulationPaper · Jan 2026 · checked 10 Oct 2026
- 9BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D ManipulationPaper · Aug 2026 · checked 10 Oct 2026
- 10What Are We Actually Benchmarking in Robot Manipulation? (2026 audit)Paper · Jun 2026 · checked 10 Oct 2026
- 11RLBench LICENSE file (Imperial College London RLBench Licence Agreement)Repository · 2019 · checked 10 Oct 2026
- 12GitHub API: stepjam/RLBench (stars, forks, open issues, licence 'NOASSERTION')Index · 10 Oct 2026 · checked 10 Oct 2026
- 13Act3D: 3D Feature Field Transformers for Multi-Task Robotic ManipulationPaper · Jun 2023 · checked 10 Oct 2026
- 143D Diffuser Actor: Policy Diffusion with 3D Scene Representations, v1Paper · Feb 2024 · checked 10 Oct 2026
- 15SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied ManipulationPaper · May 2024 · checked 10 Oct 2026
- 16Autoregressive Action Sequence Learning for Robotic Manipulation (ARP; RA-L 2025, Appendix Table A4)Paper · Oct 2024 · checked 10 Oct 2026
- 17THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (v2)Paper · Feb 2024 · checked 10 Oct 2026
- 18PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
- 19A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA EvaluationPaper · Jun 2026 · checked 10 Oct 2026
- 20CoppeliaSim website (editions and Edu terms)Official site · 2026 · checked 10 Oct 2026
- 21GitHub API: forks MohitShridhar/RLBench and markusgrotz/RLBench (licence field)Index · 10 Oct 2026 · checked 10 Oct 2026
- 22RLBench commit historyRepository · 25 Jan 2025 · checked 10 Oct 2026
- 23PerAct repository README (pre-generated RLBench datasets, licences)Repository · May 2024 · checked 11 Oct 2026
- 24Instruction-driven history-aware policies for robotic manipulations (Hiveformer; 74 RLBench tasks)Paper · Sep 2022 · checked 11 Oct 2026
- 25RLBench: The Robot Learning Benchmark & Learning Environment (arXiv abstract page)Paper · Sep 2019 · checked 10 Oct 2026
- 26Crossref record for DOI 10.1109/LRA.2020.2974707 (RA-L vol. 5, no. 2, pp. 3019-3026)Index · Apr 2020 · checked 11 Oct 2026
- 27rlbench/__init__.py (__version__ = '1.2.0')Repository · 2024 · checked 10 Oct 2026
- 28RLBench releases (latest tag 1.1.0, 2021-05-31)Repository · 31 May 2021 · checked 10 Oct 2026
- 29Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy (GemBench)Paper · Oct 2024 · checked 11 Oct 2026
- 30rlbench/tasks directory and tasks/__init__.py (task files; FS and MT task sets)Repository · 2024 · checked 10 Oct 2026
- 31RLBench issue #268: Bulb holder not visible in demonstrations generated for light_bulb_in (open)Repository · Jan 2025 · checked 11 Oct 2026
- 32PyRep repository (MIT licence)Repository · Aug 2024 · checked 10 Oct 2026
- 33RLBench project siteOfficial site · 2020 · checked 10 Oct 2026
- 34RLBench issue #171: Baseline Imitation Learning Policies & ResultsRepository · Jun 2022 · checked 11 Oct 2026
Where we searched for missing information
sim_to_real: RLBench paper (no real-robot experiment); README and site; PerAct, RVT-2, BridgeVLA (real results on separate tasks); THE COLOSSEUM (perturbation comparison, one method); PolaRiS and 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (related work only). The shared web-search budget ran out before a dedicated search for RLBench sim-real studies.
leaderboard: Project site, README, issue #171, two web searches for RLBench-18 results (2026-10-10).
license_assets: README acknowledgements, LICENSE, repository root.
prior claim 'forks mislabel MIT': GitHub API licence fields for stepjam/RLBench, MohitShridhar/RLBench (PerAct fork), markusgrotz/RLBench (PerAct2 fork); PyPI for 'rlbench' and 'pyrep'.
Change history
- Created at full depth from the checked basic entry and primary sources. Added the RLBench-18 score series, licence clause details, the CoppeliaSim Edu terms, derived benchmarks and reporting issues. Checking continued into 2026-10-11 (local time); sources opened then carry that access date.
- Published as a full entry.