BridgeData V2
BridgeData V2: A Dataset for Robot Learning at Scale
BridgeData V2 is a set of 60,096 robot trajectories (recorded robot runs) on a low-cost arm, mostly in toy kitchens. It is used as training data, and its robot setup is reused for tests.123
What a score here does not tell you Inferred
- How one paper's numbers compare with another paper's numbers.Each paper rebuilds its own toy-kitchen tasks.
- Whether a low error on recorded data means success on the real robot.The validation error of a policy (the robot's control model) correlated negatively with its real success. Validation error measures how far the policy's actions are from the recorded ones.
- How a policy handles work outside toy kitchens.The data comes from one lab and one low-cost arm, and the tasks need little precision.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| SimplerEnv (simulated WidowX + Bridge) May 2024 | Pearson correlation r = 0.827, 0.575, 1.000 and 0.990 for success on the 4 tasks. MMRV (a measure of how often two rankings disagree) was 0 on 3 tasks and 0.111 on one.15 | The same 3 policies (RT-1-X, Octo-Base and Octo-Small) ran 4 Bridge tasks in simulation and on a real WidowX arm, with 24 real trials per task. The study’s authors described the result as “strong correlation”. | The benchmark’s authors |
| Offline action error on Bridge data (SimplerEnv baseline) May 2024 | Pearson r = -0.951, -0.342, -0.857 and -1.000 on the 4 tasks. MMRV was 0.389, 0.194, 0.125 and 0.366.15 | The same 3 policies were ranked by action error on 25 Bridge validation trajectories, and the ranking was compared with real success on 4 tasks. The study’s authors described the result as “not a good proxy”. | The benchmark’s authors |
| AutoEval re-test of SIMPLER Bridge scenes Mar 2025 | Pearson r = 0.548 as the mean of 4 tasks, with a range of 0.131 to 0.942. MMRV was 0.207.16 | The same 6 policies ran 4 Bridge tasks in SIMPLER and in real trials run by humans, with 50 rollouts (test runs) each. The study’s authors described the result as “policy dependent”. | The benchmark’s authors |
| Offline action error on Bridge data (AutoEval baseline) Mar 2025 | Pearson r of about -0.26, as the mean of 5 tasks16 | Policies were ranked by action error on 400 Bridge validation trajectories, and the ranking was compared with real success in human-run trials on 5 tasks. The study’s authors described the result as “negatively correlates”. | The benchmark’s authors |
| WorldGym (world model) May 2025 | Pearson r = 0.78 for per-task success. The policies' mean scores were within 3.3 points of the real ones on average.17 | RT-1-X, Octo and OpenVLA were run in a world model (a model that predicts what the camera will see next) from the first frames of OpenVLA's 170 real Bridge trials on 17 tasks. The results were compared task by task. The study’s authors described the result as “highly correlate”. | An independent group |
Benchmarks built on BridgeData V2: SimplerEnv (WidowX), AutoEval, WorldGym, RobotArena ∞ BridgeSim.151617+1
Our assessment Opinion
Bridge results from different papers are not comparable.
Reasoning
'Tested on Bridge' names a robot setup. It does not name a fixed test. Two papers' Bridge numbers are only comparable if they used the same tasks, objects and trial counts, which is rare.
Confidence: high
The setup is easy to reproduce, but the tasks are simple.
Reasoning
The low-cost, simple setup is the reason it became a shared reference. It is also its limit, because toy kitchens and low-precision tasks say little about harder real-world work.
Confidence: medium
Treat simulated Bridge scores as weak evidence about real performance.
Reasoning
Simulated Bridge scores (SimplerEnv) are weak evidence about real performance. The paired studies are small and they disagree. The simulated test can also be matched by training close to it.
Confidence: medium
Known problems 8, 1 disputed
Several copies exist with different sizes and names
The released copies hold from 28,935 to 60,064 episodes and have different names.11920+4
Details
The paper and site give 60,096 trajectories; the PMLR abstract gives 53,896. Released copies: Berkeley RLDS bridge_dataset 1.0.0 with 60,064 episodes (renamed bridge_orig by OpenVLA), the OXE copy 'bridge' with 28,935 episodes, and an undocumented Google copy bridge_data_v2 0.0.1 with 60,063. Papers that say they trained on 'Bridge' may mean any of these.
A policy trained beside the simulated Bridge test can match top scores
A small policy with 22 million parameters, trained on data recorded beside the simulated test, nearly matches the best score.25
Details
SimplerEnv's WidowX test is meant to be passed by policies trained on real BridgeData V2, but it does not restrict training data. A 2026 audit trained a 22M-parameter policy per task on 120 scripted demonstrations recorded in simulation beside the test and reached 94.8% (91/96), against 95.8% (92/96) for X-VLA. The audit says a high score alone is not evidence of capability on this test.
There is no standard real-robot test
Papers use different tasks and rebuilt scenes, so their scores cannot be compared directly.453+4
Details
Each paper builds its own Bridge-setup tasks: the Bridge paper 14 tasks, OXE its own, OpenVLA 17, SpatialVLA 7 suites, X-VLA 5, SimplerEnv 4, AutoEval 5. OpenVLA reproduced the sink scene with 'rough approximations' and could not buy the original objects. RT-1-X scored 27% in OXE's Bridge test, 18.5% in OpenVLA's, and 0 to 4% success on SimplerEnv's real tasks.
Simulated proxies disagree with each otherDisputed
Studies disagree on how well simulated Bridge scenes, used as stand-ins for real tests, track real results.151618
Other view: SimplerEnv's own study found near-perfect ranking agreement with real results on 3 of 4 tasks.15
Details
SimplerEnv's own study found close agreement with real WidowX results for 3 policies. AutoEval, with 6 policies, found SIMPLER's accuracy depends on the policy (mean r 0.548 by our calculation). RobotArena ∞ found all tested policies score much higher on SIMPLER's 4 scenes than on its 70 Bridge-derived scenes and says SIMPLER may overestimate performance; it has no real-robot comparison.
SimplerEnv reports MMRV 0 on 3 of 4 Bridge tasks and Pearson r up to 1.000 for its 3 policies.
Offline action error is a poor predictor of real success
Lower validation error did not mean higher real success.1516
Details
Ranking policies by action error on Bridge validation data gave negative correlations with real success in SimplerEnv (r -0.342 to -1.000 on 4 tasks) and AutoEval (mean about -0.26).
Every demonstration starts with an all-zero action
Policies trained without removing the all-zero first action froze.35
Details
OpenVLA found that each demonstration records an all-zero action at the first step. Training without removing it gave policies that froze. RT-2-X was trained without this filtering; OpenVLA says the RT-2-X developers queried the second-most-likely action in the OXE Bridge evaluations, which the OXE paper does not mention.
Crowdsourced language labels contain errors
Some language labels do not match what the robot did.426
Details
Labels were added after collection through a crowdsourcing platform. A user listed mismatched instructions on the OXE tracker; an OXE author replied the labels were crowdsourced and 'there is some label noise', and the updated Berkeley copy has the same noise.
The authors say the data collection is narrow
The data comes from a single lab and one robot type, and the tasks are low-precision.4
Details
The authors list as limits that tasks are 'generally low-precision', that data comes from a single institution, and that others may find it hard to standardise on the same robot.
Details
About
- What it is
- Dataset42
More
Presented as a dataset with training code and checkpoints. No fixed test set or scoring rule.
- Released
- August 2023, at CoRL 2023119
More
arXiv v1 2023-08-24. Published at CoRL 2023.
Raw data files on the Berkeley server are dated 2023-06-20 (scripted) and 2023-08-20 (demonstrations). arXiv v3: 2024-01-17.
- Version
- Raw zip files and an RLDS 1.0.0 copy. Several other copies exist.272823
More
Raw: demos_8_17.zip (411 GB) and scripted_6_18.zip (30 GB). Pre-processed RLDS copy bridge_dataset 1.0.0 at 256x256 (about 124 GB). No repo tags. Other copies carry other names and counts (items).
Berkeley RLDS bridge_dataset 1.0.0 · 60,064 episodes: 53,192 train, 6,872 val (2023-09-21). OpenVLA renames it bridge_orig.2023
OXE copy 'bridge' 0.1.0 · 28,935 episodes: 25,460 train, 3,475 test, 387.49 GiB. An early partial upload (OXE author, 2023-12).212924
Google bucket copy bridge_data_v2 0.0.1 · 60,063 episodes: 53,191 train, 6,872 val. Created 2024-08-12; not mentioned on the site, repo or OXE spreadsheet.22
Checkpoints · Released for GCBC, D-GCBC, LCBC, GCIQL and CRL (2023-09). The README says ACT and RT-1 checkpoints are not available.3028
Setup
- Runs in
- Real robots4
More
The paper's evaluations ran on real WidowX arms. The setup is also simulated by SimplerEnv and replicated in world models (see validity).
- Robot
- One arm4
- Robot model
- WidowX 25042
More
WidowX 250 6-DoF arm; fixed over-the-shoulder RGB-D camera, two RGB cameras moved every 50 trajectories, wrist camera; VR-controller teleoperation; 5 Hz control
The paper puts the setup cost at about $4,000. Most data has only the fixed camera view.
- Tasks
- 13 skills42
More
13 skills (e.g. pick-and-place, pushing, wiping, folding, stacking, sweeping, opening doors and drawers). The paper gives no separate task count.
A skill is 'a group of trajectories that require similar motions' (Table 5). Not comparable with task counts of other datasets.
- Training data
- 60,096 trajectories142+1
More
60,096 trajectories: 50,365 teleoperated demonstrations and 9,731 scripted pick-and-place rollouts (84% human, 16% scripted). The PMLR abstract says 53,896.
CONFLICT: arXiv v1-v3 and the site say 60,096; the CoRL/PMLR abstract says 53,896. Released RLDS copies hold 60,064 (Berkeley) and 60,063 (Google) episodes.
Average length 38 steps at 5 Hz · 640x480 images; crowdsourced language labels added after collection24
All-zero first action · OpenVLA reports every demonstration starts with an all-zero action; training without removing it made policies freeze.3
- Size
- 441 GB of raw data2723
More
Raw: 411 GB demonstrations + 30 GB scripted (JPEG, PNG, pkl). RLDS at 256x256: about 124 GB.
- Changes at test
- New positions, objects, scenes and visuals4
More
Paper tests: seen tasks with new object positions, distractors and lighting; unseen objects and environments; and a second lab with a different setup.
Scoring and access
- Scored by
- Success rate4
More
Later papers sometimes give partial credit (0.5) for some tasks, e.g. OpenVLA.
- Score
- Success rate. Each paper sets its own tasks.4
More
Paper: 8 seen and 6 unseen tasks, 10 trials each, 6 methods (GCBC, D-GCBC, ACT, CRL, LCBC, RT-1). Later papers define their own Bridge-setup tasks.
Paper results · Seen tasks average: GCBC 0.49, D-GCBC 0.49, ACT 0.41, CRL 0.42, LCBC 0.23, RT-1 0.49. Unseen: 0.60, 0.55, 0.28, 0.52, 0.08, 0.50.4
Later real Bridge-setup suites · OXE: Bridge tasks at two labs. OpenVLA: 17 tasks x 10 trials. SpatialVLA: 7 task suites. X-VLA: 5 tasks. AutoEval: 5 tasks x 50 rollouts. Each defines its own tasks and objects.537+2
- Trials
- 10 per task in the paper4315+1
More
Paper: 10 trials per task. OpenVLA: 10 per task (170 per policy). SimplerEnv real: 24 per task. AutoEval: 50 per policy and task.
- Who runs it
- Each team tests its own model Inferred4316
More
Each paper runs its own trials. AutoEval, a separate project, offers public autonomous Bridge cells, but its results go back to the submitter.
- Error bars
- Sometimes reported Inferred43
More
The paper reports averages over 10 trials without error bars. OpenVLA reports standard errors for its Bridge results.
- Code licence
- MIT121314
More
LICENSE: MIT, Copyright (c) 2023 Robotic AI & Learning Lab Berkeley. The robot controller repo (bridge_data_robot) is also MIT.
- Data licence
- CC BY 4.02
More
The site states all data is provided under CC BY 4.0.
No licence file on the data server. Third-party copies on Hugging Face carry other labels (items); they do not change the original terms.
apache-2.0 · IPEC-COMMUNITY/bridge_orig_lerobot (91,062 downloads, Hub 'downloads' field)32
- Asset licence
- not applicable Inferred27
More
Real-robot recordings only; no 3D assets are distributed.
- Commercial use
- Allowed Inferred212
More
MIT code and CC BY 4.0 data both allow commercial use with attribution. Not legal advice.
- Published at
- Proceedings of the 7th Conference on Robot Learning, PMLR 229:1723-173619
Sources 32
- 1BridgeData V2: A Dataset for Robot Learning at Scale (arXiv abstract page)Paper · Aug 2023 · checked 10 Oct 2026
- 2BridgeData V2 project siteOfficial site · Aug 2023 · checked 10 Oct 2026
- 3OpenVLA (Section 5.1; Appendices B.1, C)Paper · Jun 2024 · checked 10 Oct 2026
- 4BridgeData V2 paper, full text v3Paper · Jan 2024 · checked 10 Oct 2026
- 5Open X-Embodiment paper v9 (Table I Bridge results; Table II)Paper · Oct 2023 · checked 10 Oct 2026
- 6Octo: An Open-Source Generalist Robot PolicyPaper · May 2024 · checked 10 Oct 2026
- 7SpatialVLA (real WidowX BridgeV2 evaluation)Paper · Jan 2025 · checked 10 Oct 2026
- 8X-VLA (real-world BridgeData-v2 benchmark tasks)Paper · Oct 2025 · checked 10 Oct 2026
- 9OpenVLA-OFT: Fine-Tuning Vision-Language-Action Models (Appendix, BridgeData V2)Paper · Feb 2025 · checked 10 Oct 2026
- 10π0 paper (Bridge v2 in pretraining mixture)Paper · Oct 2024 · checked 10 Oct 2026
- 11GR00T N1 paper (Bridge-v2 among OXE subsets)Paper · Mar 2025 · checked 10 Oct 2026
- 12bridge_data_v2 LICENSERepository · 2023 · checked 10 Oct 2026
- 13GitHub API: rail-berkeley/bridge_data_v2Index · 10 Oct 2026 · checked 10 Oct 2026
- 14GitHub API: rail-berkeley/bridge_data_robot (MIT)Index · 10 Oct 2026 · checked 10 Oct 2026
- 15SimplerEnv: Evaluating Real-World Robot Manipulation Policies in Simulation (Tables V, XII; Appendix B)Paper · May 2024 · checked 10 Oct 2026
- 16AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World (v2; Tables 1-4, Figure 7)Paper · Mar 2025 · checked 10 Oct 2026
- 17WorldGym: World Model as An Environment for Policy Evaluation (v3)Paper · May 2025 · checked 10 Oct 2026
- 18RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationPaper · Oct 2025 · checked 10 Oct 2026
- 19BridgeData V2, PMLR 229 (CoRL 2023) pagePaper · Dec 2023 · checked 10 Oct 2026
- 20Berkeley RLDS bridge_dataset 1.0.0 dataset_info.jsonDataset page · 21 Sep 2023 · checked 10 Oct 2026
- 21TensorFlow Datasets catalog: bridge (OXE copy)Dataset page · 10 Oct 2026 · checked 10 Oct 2026
- 22bridge_data_v2 0.0.1 dataset_info.json in gs://gresearch/roboticsDataset page · 12 Aug 2024 · checked 10 Oct 2026
- 23OpenVLA README (download Berkeley RLDS copy, 124 GB, rename to bridge_orig)Repository · Sep 2024 · checked 10 Oct 2026
- 24OXE issue #30: OXE Bridge copy is an early upload (author reply)Repository · Dec 2023 · checked 10 Oct 2026
- 25What Are We Actually Benchmarking in Robot Manipulation? (Section 6, data-source dependence)Paper · Jun 2026 · checked 10 Oct 2026
- 26OXE issue #15: Bridge label noise (author reply)Repository · Nov 2023 · checked 10 Oct 2026
- 27Berkeley data server listing (raw zips, tfds folder)Dataset page · Sep 2023 · checked 10 Oct 2026
- 28bridge_data_v2 READMERepository · Mar 2024 · checked 10 Oct 2026
- 29OXE bridge 0.1.0 dataset_info.jsonDataset page · Oct 2023 · checked 10 Oct 2026
- 30Released checkpoints listingDataset page · Sep 2023 · checked 10 Oct 2026
- 31bridge_data_v2 commit historyRepository · 17 Mar 2024 · checked 10 Oct 2026
- 32Hugging Face: IPEC-COMMUNITY/bridge_orig_lerobot (third-party copy)Secondary · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
validity / sim_to_real: Bridge paper v3 (no sim experiments); SimplerEnv Tables V, XII; AutoEval v2 Tables 1-4 and Figure 7; WorldGym v3; RobotArena ∞ (no real correlation); OXE paper; 2026 audit 2606.04233 (Section 6); IRASim, WorldEval and dWorldEval leads (WorldEval and dWorldEval do not use Bridge).
tasks, objects: Paper Sections 3.2-3.3 and Table 5; site. No task count beyond 13 skills; objects given as 'more than 100'.
top_score: No single headline score exists for the dataset; real Bridge suites differ by paper. SimplerEnv WidowX scores belong to the SimplerEnv record.
Change history
- Created at full depth from primary sources, starting from the checked basic entry and research/raw/inventory/core-real.json. Added: episode counts of every copy, all-zero first action, label noise, checkpoint gaps, five validity studies (two recomputed or read from figures), two-lab measurements with our correlation figures, 2026 audit finding on SimplerEnv WidowX, adoption. Corrected the basic entry: at Lab 2, two of six methods improved (ACT, CRL).
- Published as a full entry.