RLBench

RLBench: The Robot Learning Benchmark & Learning Environment

How to read this picture

RLBench is a set of 100 simulated tasks for one Franka Panda robot arm. Papers on robot manipulation policies (the robot's control models) use it, and most report results on an 18-task subset.12

Sources
Last checked 10 Oct 2026Full entry55 of 66 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm3
Franka Emika Panda (simulated)
Used by
Mainly papers on 3D manipulation policies245+5
1,139 citations
Licence
Not allowed1112
Commercial use: not allowed

What a score here does not tell you Inferred

  1. How well a policy will do on a real robot.No study has compared RLBench and real-robot scores for the same policies.
  2. How well a policy handles new objects.The common protocol tests only objects seen in training.
  3. How well a policy copes with changes to the scene.THE COLOSSEUM, a test built on RLBench, shows large drops in success under scene changes.
ChartPublished scores over time
95% AND ABOVE5060708090100202420252026RVT · 62.9% · 2023-06Act3D · 65% · 2023-063D Diffuser Actor · 81.3% · 2024-02SAM-E · 70.6% · 2024-05RVT-2 · 81.4% · 2024-06ARP+ · 84.9% · 2024-10SAM2Act · 86.8% · 2025-01BridgeVLA · 88.2% · 2025-06ActiveVLA · 91.8% · 2026-01BridgeVLA++ · 93.7% · 2026-08RVT 62.9%BridgeVLA++ 93.7%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.41314+7

Comparisons with real robots

Tried on real robots, not compared Inferred125+4

Details

The RLBench paper offers domain randomisation and swappable arms for sim-to-real research but has no real-robot experiment. The closest evidence is THE COLOSSEUM (RSS 2024), built on RLBench: it mirrored 4 tasks with 3D-printed objects on a real Franka, trained one PerAct on 5 real demonstrations per task and another in simulation, and tested each under matching perturbations (10 episodes × 3 runs). Per-factor R² between the two ranged from 0.46 to 0.94 (mean 0.614). These are two separately trained models of one method compared across perturbation conditions, so we do not count it as a paired evaluation of the same policies. PolaRiS and the 2026 sim-real recipe paper mention RLBench only in related work.

Our assessment Opinion

Check which RLBench subset and protocol a number uses.

Reasoning

When a paper reports 'RLBench', it almost always means PerAct's 18-task protocol, not the 100-task benchmark or its few-shot challenge. Check the task subset, the number of demonstrations and the inputs before comparing.

Confidence: high

A high score shows that a policy fits 18 known tasks. It does not show that the policy copes with scene changes.

Reasoning

A high RLBench-18 score shows a policy fits 18 known tasks with seen objects in one scene. Scene perturbations cut scores sharply, and no study has checked RLBench scores against real-robot results for the same policies.

Confidence: medium

Ignore gaps of a few points between top results.

Reasoning

Gaps of a few points at the top are not reliable. Data-loader details move the same model by several points, and papers re-report earlier models with different numbers.

Confidence: high

Commercial users need permission from Imperial College London. The default edition of the simulator also excludes commercial use.

Reasoning

Companies should read the licence before using RLBench. It and the default simulator edition exclude commercial use and product research without permission. This is our reading, not legal advice.

Confidence: high

Known problems 7

  1. The licence is non-commercial and the simulator edition is for education only

    The licence does not allow commercial use. The default simulator edition is for schools and universities only.11203+1

    Details

    RLBench's custom licence allows only non-commercial, internal or academic research, bans transfer and sub-licensing, lets Imperial terminate at any time, and asks publications to credit Imperial and send it a copy. The README installs CoppeliaSim Edu, whose terms exclude companies, research institutions and commercial use. Main forks keep the custom licence (GitHub shows NOASSERTION).

  2. Maintenance has stopped and 3D assets lack licences

    The code has had no functional changes since July 2024 and uses an old simulator version. The third-party 3D models have no licence information.22320+1

    Details

    The last functional commit is from July 2024 and the code pins CoppeliaSim 4.1.0 (the vendor's current version is 4.10.0). 94 issues are open. The README credits 3D models to five third-party model sites without per-model licence information.

  3. The same model gets different scores

    The score of one model, RVT-2, falls from 81.4% to 74.1% after one change to how training data is loaded.479+2

    Details

    PerAct's paper gives no average; RVT re-evaluated its released model at 49.4%. BridgeVLA reports 88.2%; its follow-up BridgeVLA++ reports BridgeVLA at 90.5 ± 1.1%. ARP found that RVT-2's 81.4% depends on a data-loader detail inherited from C2F-ARM: removing a randomised timestep gives 77.0%, and giving correct timesteps gives 74.1%. RVT notes that earlier baselines picked the best model per task on validation sets, which overstates their multi-task scores.

  4. Top scores are approaching 100%

    The best average over the 18 tasks is 93.7%.489

    Details

    RLBench-18 averages rose from 62.9% (RVT, 2023-06) to 91.8% (ActiveVLA, 2026-01) and 93.7% (BridgeVLA++, 2026-08). Many individual tasks are at or near 100%; in BridgeVLA++ the place cups task is at 76.8%.

  5. Results vary from one evaluation run to the next

    The motion planner makes random choices, so results vary between runs. Some papers report only one run.42317

    Details

    RVT runs each model five times on the same 25 episodes per task because the sampling-based motion planner is random. PerAct distributes fixed pre-generated datasets because data generation samples scenes randomly. Some papers report one evaluation run.

  6. Scores fall when the scene changes

    In THE COLOSSEUM, success drops by 30 to 50% for each type of scene change, and by at least 75% when changes are combined.17

    Details

    THE COLOSSEUM perturbs 20 RLBench tasks along 14 factors such as colour, texture, lighting, distractors and camera pose. It reports success dropping 30 to 50% per factor and at least 75% when factors are combined. PerAct trained without perturbations fell from 34.5% to 6.4% with all perturbations on.

  7. Most RLBench results come from a modified 18-task subset

    Most papers report results on a modified subset of 18 tasks defined by PerAct, instead of the full 100-task benchmark.2241+1

    Details

    Most recent papers report PerAct's RLBench-18 protocol: 18 tasks, some modified to add variations (249 in total), 100 demonstrations and 25 test episodes per task, four 128 × 128 RGB-D cameras and keyframe actions. Others use 74 tasks (Hiveformer) or their own subsets. The paper's own few-shot challenge is rarely reported, and the paper (10% of tasks held out) and README (5 test tasks) disagree on its split.

Details

About

What it is
Benchmark1
More

The paper presents RLBench as a benchmark and learning environment with 100 tasks and a few-shot challenge.

Built by
Imperial College London (Dyson Robotics Lab)1
More

Authors: Stephen James, Zicong Ma, David Rovick Arrojo, Andrew J. Davison (Dyson Robotics Lab and UROP, Imperial College London). The research was supported by Dyson Technology Ltd.

Released
September 2019, in IEEE Robotics and Automation Letters (RA-L) 202025263+1
More

arXiv v1 on 2019-09-26 (repository created the same day). Published in IEEE Robotics and Automation Letters 5(2), pp. 3019-3026 (April 2020), presented at ICRA 2020.

Version
1.2.0. The last release tag is 1.1.0.27283
More

1.2.0 in the code; the last GitHub release tag is 1.1.0 (2021-05-31). Not on PyPI; installed from GitHub.

RLBench-18 (PerAct protocol) · 18 tasks, some modified to add variations (249 variations in total), 100 training demos and 25 test episodes per task. Defined by PerAct (CoRL 2022), not by the RLBench authors. Used by most recent papers.2

RLBench-74 (Hiveformer protocol) · 74 tasks, CoRL 2022.2429

Few-shot task sets · FS10_V1, FS25_V1, FS50_V1, FS95_V1: 10 to 95 training tasks plus 5 test tasks. Multi-task sets MT15_V1 to MT100_V1.330

Last update
January 2025. The last commit fixed a typo.223
More

Last commit 2025-01-25 (typo fix in a tutorial). Last functional changes 2024-07-02/03 (Gymnasium support, arm velocity and acceleration limits).

README announcements: v1.2.0 on 2022-02-18 (breaking action-mode API changes); shaped rewards for two tasks on 2022-05-11.

Status
No functional changes since July 2024. Still used in new papers. Inferred221231
More

No functional change since July 2024; last commit January 2025. Still used in new papers.

94 open issues. Issue #268 (2025-01, open): the bulb holder is missing from generated light_bulb_in demonstrations; the reporter switched tasks.

Setup

Runs in
Simulation1
Simulator
CoppeliaSim 4.1, through PyRep3132
More

The README downloads CoppeliaSim Edu 4.1.0 for Ubuntu 20.04. The CoppeliaSim website lists V4.10.0 as current. PyRep is installed from GitHub; the PyPI name 'pyrep' belongs to an unrelated AGPL package.

Robot
One arm3
Robot model
Franka Emika Panda (simulated)31
More

Mico, Jaco, Sawyer and UR5 can be swapped in, but the README says the arm should remain the Franka Panda for benchmarking.

Setting
Tabletop1
Tasks
100 tasks, of which 18 are commonly used130
More

100 tasks in the paper; 106 task files in the repository today

Repository count by us: .py files in rlbench/tasks excluding __init__.py (2026-10-10). Recent papers mostly use the 18-task subset.

Scenes
1 scene1
More

1 shared table scene

Training data
Unlimited, generated by a motion planner123
More

Unlimited demonstrations generated by a motion planner (OMPL) from task waypoints; lengths 100 to 1,000 steps. No fixed official dataset.

PerAct pre-generated data · Train 100, validation 25 and test 25 episodes per task for the 18 tasks, about 116 GB on Google Drive. PerAct says using them helps reproducibility because data generation samples scenes randomly.23

Changes at test
New object positions. The objects are the same as in training.21
More

In the common RLBench-18 protocol, test episodes use new object poses and sampled variations (colours, sizes, targets) that also appear in training. Train and test objects are the same.

PerAct states that it does not test generalisation to unseen objects. The original few-shot challenge (new tasks) is rarely reported; derived benchmarks such as GemBench and AGNOSTOS test new tasks.

Scoring and access

Scored by
Success rate12
More

Each task has a sparse reward of +1 on completion, checked by task-specific success conditions.

Score
Success rate over 18 tasks21
More

In the RLBench-18 protocol: one multi-task policy, each test episode scored 0 or 100, averaged over 25 episodes per task and over the 18 tasks. The paper's own challenge asks for 1-, 5- and 20-shot success on held-out tasks.

Trials
25 per task, often repeated in 5 runs249
More

PerAct · 25 episodes per task (450 in total), at most 25 keyframe steps; best checkpoint chosen on a separate 25-episode validation set.2

RVT and later · Each model run 5 times on the same 25 episodes per task, because the sampling-based motion planner is random; mean and standard deviation reported.4

BridgeVLA++ · Mean ± std over five random seeds, 25 evaluation episodes per seed.9

Who runs it
Each team tests its own model Inferred33334
More

No submission process or organiser-run evaluation on the site or README.

Error bars
Sometimes reported Inferred246+2
More

RVT, RVT-2, 3D Diffuser Actor, SAM2Act, BridgeVLA and BridgeVLA++ report standard deviations over 5 runs. PerAct reports one evaluation; COLOSSEUM reports one training seed and one evaluation seed.

Leaderboard
None. Scores are only in papers.33334
More

No leaderboard on the project site or README. In issue #171 (2022) the maintainer said no official baselines exist because simple policies did not work on these sparse-reward tasks.

Code licence
Custom licence, non-commercial use only1112
More

Custom licence: free, non-exclusive, non-transferable; use only for non-commercial, internal or academic research. BSD terms for BSD elements.

Key clauses: 2(b) bans commercial use, including research to develop products for sale and paid services; commercial use requires contacting researchcontracts.engineering@imperial.ac.uk. 2(d) bans transfer and sub-licensing. 2(f) requires publications to credit the software as licensed from Imperial and to send Imperial a copy. 7(a) Imperial may terminate at any time; 7(d) restrictions expire 10 years after first use. English law. GitHub shows 'NOASSERTION'.

Data licence
not-applicable Inferred123
More

No official dataset is distributed; users generate demonstrations. PerAct's pre-generated data are produced with RLBench and hosted by the PerAct authors.

Asset licence
Unknown
More

The README says 3D models came from turbosquid.com, cgtrader.com, free3d.com, thingiverse.com and cadnav.com. No per-model licence or attribution file was found in the README or LICENSE; the RLBench licence covers 'the Software'.

Access
Open. Downloading the code binds the user to the licence.31120
More

Code on GitHub with no registration; downloading binds the user to the licence. The simulator download requires accepting CoppeliaSim's terms.

Commercial use
Not allowed Inferred1120
More

From clause 2(b) of the RLBench licence, plus the CoppeliaSim Edu terms. Not legal advice.

Published at
IEEE Robotics and Automation Letters, 2020263
More

DOI 10.1109/LRA.2020.2974707. README announcement 2020-01-28: accepted to RA-L with presentation at ICRA.

Sources 34

  1. 1RLBench paper, full text v1Paper · Sep 2019 · checked 10 Oct 2026
  2. 2Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation (PerAct)Paper · Sep 2022 · checked 10 Oct 2026
  3. 3RLBench GitHub README (install, task sets, announcements, acknowledgements)Repository · 2024 · checked 10 Oct 2026
  4. 4RVT: Robotic View Transformer for 3D Object ManipulationPaper · Jun 2023 · checked 10 Oct 2026
  5. 5RVT-2: Learning Precise Manipulation from Few DemonstrationsPaper · Jun 2024 · checked 10 Oct 2026
  6. 6SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic ManipulationPaper · Jan 2025 · checked 10 Oct 2026
  7. 7BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language ModelsPaper · Jun 2025 · checked 10 Oct 2026
  8. 8ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic ManipulationPaper · Jan 2026 · checked 10 Oct 2026
  9. 9BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D ManipulationPaper · Aug 2026 · checked 10 Oct 2026
  10. 10What Are We Actually Benchmarking in Robot Manipulation? (2026 audit)Paper · Jun 2026 · checked 10 Oct 2026
  11. 11RLBench LICENSE file (Imperial College London RLBench Licence Agreement)Repository · 2019 · checked 10 Oct 2026
  12. 12GitHub API: stepjam/RLBench (stars, forks, open issues, licence 'NOASSERTION')Index · 10 Oct 2026 · checked 10 Oct 2026
  13. 13Act3D: 3D Feature Field Transformers for Multi-Task Robotic ManipulationPaper · Jun 2023 · checked 10 Oct 2026
  14. 143D Diffuser Actor: Policy Diffusion with 3D Scene Representations, v1Paper · Feb 2024 · checked 10 Oct 2026
  15. 15SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied ManipulationPaper · May 2024 · checked 10 Oct 2026
  16. 16Autoregressive Action Sequence Learning for Robotic Manipulation (ARP; RA-L 2025, Appendix Table A4)Paper · Oct 2024 · checked 10 Oct 2026
  17. 17THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (v2)Paper · Feb 2024 · checked 10 Oct 2026
  18. 18PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
  19. 19A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA EvaluationPaper · Jun 2026 · checked 10 Oct 2026
  20. 20CoppeliaSim website (editions and Edu terms)Official site · 2026 · checked 10 Oct 2026
  21. 21GitHub API: forks MohitShridhar/RLBench and markusgrotz/RLBench (licence field)Index · 10 Oct 2026 · checked 10 Oct 2026
  22. 22RLBench commit historyRepository · 25 Jan 2025 · checked 10 Oct 2026
  23. 23PerAct repository README (pre-generated RLBench datasets, licences)Repository · May 2024 · checked 11 Oct 2026
  24. 24Instruction-driven history-aware policies for robotic manipulations (Hiveformer; 74 RLBench tasks)Paper · Sep 2022 · checked 11 Oct 2026
  25. 25RLBench: The Robot Learning Benchmark & Learning Environment (arXiv abstract page)Paper · Sep 2019 · checked 10 Oct 2026
  26. 26Crossref record for DOI 10.1109/LRA.2020.2974707 (RA-L vol. 5, no. 2, pp. 3019-3026)Index · Apr 2020 · checked 11 Oct 2026
  27. 27rlbench/__init__.py (__version__ = '1.2.0')Repository · 2024 · checked 10 Oct 2026
  28. 28RLBench releases (latest tag 1.1.0, 2021-05-31)Repository · 31 May 2021 · checked 10 Oct 2026
  29. 29Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy (GemBench)Paper · Oct 2024 · checked 11 Oct 2026
  30. 30rlbench/tasks directory and tasks/__init__.py (task files; FS and MT task sets)Repository · 2024 · checked 10 Oct 2026
  31. 31RLBench issue #268: Bulb holder not visible in demonstrations generated for light_bulb_in (open)Repository · Jan 2025 · checked 11 Oct 2026
  32. 32PyRep repository (MIT licence)Repository · Aug 2024 · checked 10 Oct 2026
  33. 33RLBench project siteOfficial site · 2020 · checked 10 Oct 2026
  34. 34RLBench issue #171: Baseline Imitation Learning Policies & ResultsRepository · Jun 2022 · checked 11 Oct 2026
Where we searched for missing information

sim_to_real: RLBench paper (no real-robot experiment); README and site; PerAct, RVT-2, BridgeVLA (real results on separate tasks); THE COLOSSEUM (perturbation comparison, one method); PolaRiS and 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (related work only). The shared web-search budget ran out before a dedicated search for RLBench sim-real studies.

leaderboard: Project site, README, issue #171, two web searches for RLBench-18 results (2026-10-10).

license_assets: README acknowledgements, LICENSE, repository root.

prior claim 'forks mislabel MIT': GitHub API licence fields for stepjam/RLBench, MohitShridhar/RLBench (PerAct fork), markusgrotz/RLBench (PerAct2 fork); PyPI for 'rlbench' and 'pyrep'.

Change history

  1. Created at full depth from the checked basic entry and primary sources. Added the RLBench-18 score series, licence clause details, the CoppeliaSim Edu terms, derived benchmarks and reporting issues. Checking continued into 2026-10-11 (local time); sources opened then carry that access date.
  2. Published as a full entry.