Open X-Embodiment
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment (OXE) pools robot data from dozens of labs into one format, and model builders use it as training data. It has no test of its own.123
What a score here does not tell you Inferred
- How often a policy (the robot's control model) will succeed on a real robot.Some papers rank policies by action error on OXE episodes, which measures how far predicted actions are from the recorded ones. This ranked policies poorly compared with real-robot success.
- Whether results from different labs can be compared.In the RT-X study, each lab tested the models on its own tasks.
- How much of the data comes from real robots.Of the listed episodes, 38% come from simulation.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| SimplerEnv (simulated Google Robot setup) May 2024 | Pearson r = 0.924 and MMRV = 0.056 for Visual Matching. Pearson r = 0.778 and MMRV = 0.143 for Variant Aggregation. MMRV measures how often two rankings disagree.3 | The same 6 policies (3 RT-1 checkpoints, RT-1-X, RT-2-X and Octo-Base) were scored in a simulated copy of the Google Robot setup and on the real robot, over 3 task groups. The study’s authors described the result as “strong correlation”. | The benchmark’s authors |
| Offline action error on Google Robot data (SimplerEnv baseline) May 2024 | Mean over 3 tasks: Pearson r = 0.308, MMRV = 0.3753 | The same 6 policies were ranked by their action prediction error on 25 Google Robot (RT-1) training episodes. The ranking was compared with real-robot success over 3 task groups. The study’s authors described the result as “not a good proxy”. | The benchmark’s authors |
| WorldGym (world model trained on OXE) May 2025 | Pearson r = 0.78 for success on each task. The policies' mean scores in the world model differed from their real-robot means by 3.3 points on average.13 | RT-1-X, Octo and OpenVLA were run in a world model (a learned simulator) trained on 9 OXE datasets. Each run started from the first frames of one of OpenVLA's 170 real Bridge trials (17 tasks), and results were compared task by task. The study’s authors described the result as “highly correlate”. | An independent group |
Benchmarks built on Open X-Embodiment: SimplerEnv, WorldGym, RobotArena ∞.31314
Our assessment Opinion
OXE is training data. When a paper reports a score on OXE, check which test it came from.
Reasoning
OXE is training data and a reference for robot setups. It is not a test. A claim to 'evaluate on OXE' can mean real trials in an OXE lab setup, a simulated copy such as SimplerEnv, or action error on held-out episodes. Only the first two have been checked against real robots, and the third ranked policies poorly in every study we found.
Confidence: high
Check which OXE datasets a model actually used.
Reasoning
Treat OXE totals with care. The episode count mixes real and simulated data, the Bridge copy is partial, and licence terms vary by dataset. When a model report says it trained on OXE, check which datasets and versions it used.
Confidence: high
The RT-X results come from one study that others cannot repeat.
Reasoning
The RT-X study is evidence that training on data from many robots can help a given robot. It is one study, and others cannot repeat it, because RT-2-X is closed, each lab used its own tasks, and one evaluation detail went unreported. Later open models trained on OXE, such as Octo and OpenVLA, give evidence that others can check.
Confidence: medium
Known problems 8
Over a third of listed episodes are simulated
Of the listed episodes, 38% come from simulation. Most of them were added in v1.1. Inferred11516+1
Details
The paper describes '1M+ real robot trajectories'. The spreadsheet now lists 2,419,193 episodes, of which 924,553 (38.2%) come from simulated datasets: VIMA 660,103 and SPOC 233,000 (simulation per their own papers), ManiSkill 30,000, USC Cloth Sim 1,000 and Plex RoboSuite 450. Most of this arrived with v1.1, where simulated data is 88.0% of the batch. Totals that mix v1.0 and v1.1 therefore include a large share of simulated data.
Counts differ between the paper, blog and spreadsheet
The paper, the blog and the spreadsheet give different numbers of institutions, robots, datasets and episodes.18119+3
Details
Institutions: 21 (abstract), 33 academic labs (blog), 34 labs (Section III-A), 46 affiliation entries (PDF). Robots: 22 embodiments (paper) vs 27 robot names (spreadsheet). Datasets: 60 (paper) vs 72 (spreadsheet), with two more in the bucket that the spreadsheet does not list. Episodes: 1M+ (paper) vs 2,419,193 (spreadsheet). Skills and tasks: 527 and 160,266 (paper) vs more than 500 and 150,000 (blog).
The OXE copy of BridgeData V2 is an early partial upload
OXE's copy of BridgeData V2 has 28,935 of about 60,000 episodes. Some users, such as Octo and OpenVLA, train on other copies.222315+3
Details
OXE's 'bridge' dataset has 25,460 train and 3,475 test episodes (28,935 in total), against 60,096 trajectories in BridgeData V2. In December 2023 an OXE author wrote that it was uploaded at an early stage and would be updated. On 2026-10-10 the spreadsheet still lists 25,460 episodes. Octo and OpenVLA train on the Berkeley copy instead; Octo's config notes it is not the official OXE copy. A full copy (bridge_data_v2/0.0.1, 60,063 episodes) has sat in the same bucket since 2024-08-12 without documentation.
Several datasets have reported defects and one download path is half empty
Users have reported defects in several datasets. One download path given in the README has data in only 35 of its 72 dataset folders.272829+2
Details
Users report defects in individual datasets on the issue tracker. Roboturk episodes never mark their end; an author confirmed this in 2023-11 and promised a fix, and the issue is still open. Other open reports: Saytap images are all zeros (2024-07), berkeley_rpt lacks two of three camera views (2024-10), and fractal20220817_data contains failed episodes (2024-02). The README's fallback download path, gs://gdm-robotics-open-x-embodiment, has data in only 35 of its 72 dataset folders; the other 37, including bridge and fractal20220817_data, hold only folder markers (our listing, 2026-10-10; issue #104, open since 2025-07).
Licence terms of the pooled datasets are unclear
One CC BY 4.0 notice covers all the datasets. Some datasets state different terms at their sources, and some state no data licence. Inferred111520+3
Details
The README puts all non-software materials under CC BY 4.0 and asks users to cite each contributed dataset. The spreadsheet has no licence column, and the only licence file in the bucket is DROID's. At their own sources, several datasets state CC BY 4.0; ManiSkill2 states CC BY-NC 4.0 for its assets; Language Table and SPOC state only code licences.
The RT-X results cannot be re-run by others
The RT-2-X weights were not released, and each lab tested on its own tasks. The paper also does not mention a workaround in how RT-2-X's actions were chosen in the Bridge tests.35361+1
Details
RT-2-X weights were not released: an author replied in 2023-11 that there was 'no possibility to release' them, and in 2024-03 that there were no plans to release RT-2 code. Each lab evaluated on its own tasks; the text gives no per-lab trial counts or error bars. OpenVLA reports that RT-2-X froze on Bridge because every Bridge demonstration starts with an all-zero action, and that the RT-2-X developers queried the second-most-likely action in the OXE Bridge evaluations. The OXE paper does not mention this. The Table I caption also contradicts its own numbers for RT-1-X on Bridge.
Actions are only coarsely aligned across datasets
Actions are recorded in different coordinate frames and with different meanings across datasets. Model builders therefore choose subsets of the data by hand.145+1
Details
The paper converts each dataset to a 7-dimensional end-effector action but does not align coordinate frames and allows absolute or relative values, so 'the same action vector may induce very different motions for different robots'. Model builders pick subsets and weights by hand: Octo uses 25 datasets, π0 an 'OXE Magic Soup' subset. OpenVLA filtered Bridge's all-zero first actions and dropped DROID for the last third of training.
Offline action error on OXE data does not predict real success
When policies were ranked by action error on held-out OXE data, the ranking matched real-robot results poorly. In some tests the order was reversed.3
Details
SimplerEnv calls action error on held-out demonstrations a widely adopted way to select policies. On the Google Robot setup it ranked 6 policies poorly (mean Pearson r 0.308, MMRV 0.375). On Bridge tasks its correlation with real success was negative (r -0.342 to -1.000). AutoEval (Bridge) and PolaRiS (DROID) also report negative correlations; see those records.
Details
About
- What it is
- Dataset111
More
The paper and README present a dataset (the 'Open X-Embodiment Repository') and model checkpoints. There is no fixed test set, scoring rule or leaderboard.
- Built by
- Open X-Embodiment Collaboration, Google DeepMind11911+1
More
Contributors can still enrol datasets through a form linked on the project site.
Open X-Embodiment Collaboration · 293 authors on arXiv v9 (Crossref lists 279). The PDF author footnote has 46 numbered affiliations; UT Austin appears twice (entries 30 and 45).11937
Google DeepMind · Owns the GitHub repository and hosts the data bucket. README copyright: DeepMind Technologies Limited. Announced the release on its blog on 2023-10-03.112120
Institution count differs by source · 21 institutions (abstract), 33 academic labs (Google DeepMind blog), 34 labs (paper Section III-A), 46 affiliation entries (PDF footnote).18119+1
- Released
- October 2023, published at ICRA 2024211837
More
Announced 2023-10-03 on the Google DeepMind blog. arXiv v1 2023-10-13. Published at ICRA 2024.
Bucket folder markers for v1.0 datasets date from 2023-09-27 to 2023-10-05. The GitHub repo was created 2023-10-20.
- Version
- v1.0 and v1.1. The repository has no tagged releases.152012
More
The official spreadsheet groups 72 datasets into v1.0 (60) and v1.1 (12). Each dataset has its own TFDS version folder (e.g. 0.1.0). The repo has no tags or releases.
v1.0 · 60 datasets, 1,403,481 episodes, 4,435.41 GB15
v1.1 · 12 datasets (DROID, ConqHose, DobbE, FMB, IO-AI Office PicknPlace, MimicPlay, MobileALOHA, RoboSet, TidyBot, VIMA, SPOC, Plex RoboSuite), 1,015,712 episodes, 4,529.53 GB. Their bucket folders were created between 2024-03-21 and 2024-08-05.1520
Unlisted additions · bridge_data_msr (822 WidowX episodes from Microsoft Research, folder created 2024-04-11) and robo_ai_u_r5e (443 UR5e episodes from Satakunta University of Applied Sciences, 2026-04-27) are in the bucket but not in the spreadsheet.203839+1
- Last update
- April 2026. One dataset was added but not listed.392040+1
More
2026-04-27: a UR5e dataset (robo_ai_u_r5e, 443 episodes) appeared in the data bucket. It is not in the spreadsheet or README. Last repo commit 2025-11-05 (an install fix). Last paper version: arXiv v9, 2025-05-14.
Dates are bucket object timestamps and commit dates. We found no announcement for the 2026 addition.
- Status
- Maintained, with low activity Inferred402012+1
More
Low activity. One unlisted dataset was added to the bucket in April 2026. The spreadsheet and README do not list it. Reported data defects remain open.
56 open issues and pull requests (GitHub API). The last merged change is the 2025-11-05 install fix. Use as pretraining data is high (facts.used_by).
Setup
- Runs in
- Real robots1
More
Describes the RT-X evaluation. The data is also used offline (action prediction error) and through simulated or learned replicas (see validity).
- Simulator
- None for the RT-X tests, which ran on real robots. Five pooled datasets are simulation data (facts.demonstrations).115
- Robot
- Many types1
- Robot model
- 22 robot types115
More
RT-X test robots (6) · WidowX (Stanford IRIS, Berkeley RAIL), Google Robot (Google), Franka (Freiburg; Berkeley cable routing), Jaco 2 (USC), Hello Stretch (NYU), UR5 (Berkeley AUTOLab)1215
Spreadsheet morphology counts · 54 single arm, 10 mobile manipulator, 3 wheeled robot, 2 bimanual, 1 quadruped, 1 human, 1 mixed robot and human15
- Setting
- Tabletop, Kitchen, Whole home, Mixed Inferred15
More
Counted by us from the spreadsheet; one dataset can carry several tags.
- Tasks
- 160,266 tasks in 527 skills1821
More
160,266 tasks grouped into 527 skills (paper abstract). The blog says more than 500 skills and 150,000 tasks.
Skills and objects were extracted from the language annotations with the PaLM language model (Section III-B), so datasets without annotations do not contribute. Not comparable with task counts of fixed benchmarks.
- Scenes
- About 300 scenes in total, by the DROID authors' count Secondary41
More
OXE's own documents give no total. The paper plots distinct scenes per robot type (Fig. 1b) without a total. The DROID paper says OXE spans 'approximately 300 scenes total'.
- Training data
- 2.4 million episodes in 72 datasets151
More
2,419,193 episodes in 72 datasets (spreadsheet header, 2026-10-10). The paper says 1M+ trajectories from 60 datasets.
1M+ (paper) · '1M+ real robot trajectories' from 60 pooled datasets (Section III-A)1
924,553 simulated episodes (38.2%) · VIMA 660,103 and SPOC 233,000 (both simulation per their own papers), ManiSkill 30,000, USC Cloth Sim 1,000 and Plex RoboSuite 450. They make up 88.0% of the v1.1 batch.151617+1
Four datasets hold 79.2% · VIMA 27.3%, QT-Opt 24.0%, Language Table 18.3% and SPOC 9.6% of all listed episodes15
Human-body data · IO-AI Office PicknPlace (3,847 episodes) is recorded on human hands with motion capture. RoboVQA (61,153) mixes robot and human episodes.15
27 of 72 without language · 26 datasets list 'None' for language annotations and 1 states it has none. 15 were collected by scripted policies and 13 by expert policies. 26 are flagged as containing suboptimal data.15
Subsets used for training · RT-X robotics mixture: 12 datasets from 9 manipulators (paper). Octo: 800k episodes from 25 datasets. OpenVLA: 970k episodes. Octo and OpenVLA put the RT-X subset at 350K episodes; the OXE paper gives no count.145
- Size
- About 9 TB15
More
8,964.94 GB current download size for 72 datasets (spreadsheet header): v1.0 4,435.41 GB, v1.1 4,529.53 GB
The per-version split is summed by us from the rows.
- Changes at test
- New skills, objects and scenes in the RT-X study1
More
In the RT-X study: skills seen only in another robot's data, and unseen objects, backgrounds and environments. The small-data domains were tested in distribution.
Describes the paper's evaluation, Table II ('Emergent Skills' and 'RT-2 Generalization' columns). The dataset itself defines no test conditions.
Scoring and access
- Scored by
- Success rate1
More
RT-X results are success rates on tasks defined by each lab. Offline use scores action prediction error (MSE) on held-out episodes, which has no taxonomy value.
- Score
- Success rate on each lab's own tasks1
More
Success rate per lab-defined task. Small-data domains (5 datasets): RT-1-X against each dataset's 'Original Method' and an RT-1 trained on that dataset alone. Large-data domains: Bridge (WidowX) and RT-1 (Google Robot). Emergent-skill and generalisation tests on the Google Robot.
RT-1-X, small-data domains · Mean success 50% higher (relative) than the Original Method or RT-1; better than the Original Method on 4 of 5 datasets121
Large-data domains (Table I) · Bridge at Stanford IRIS / Berkeley RAIL: LCBC 13% / 13%, RT-1 40% / 30%, RT-1-X 27% / 27%, RT-2-X (55B) 50% / 30%. Google Robot: RT-1 92%, RT-1-X 73%, RT-2-X 91%.1
RT-2-X emergent skills (Table II) · RT-2-X (55B) 75.8% vs RT-2 27.3% (about 3x); 42.8% when Bridge data is removed from training. RT-2 generalisation test: 61% vs 62%.1
- Trials
- 3,600 real-robot trials in total1
- Who runs it
- Each team tests its own model Inferred12
More
Each lab ran RT-X trials on its own setup; later papers run their own trials. No organiser runs submissions.
- Error bars
- Sometimes reported Inferred15
More
The OXE paper text reports no error bars or intervals. Some later real-robot reports in OXE setups do: OpenVLA gives standard errors.
- Leaderboard
- None. Scores are only in papers. Inferred21115
More
No leaderboard on the project site, README or spreadsheet.
- Code licence
- Apache-2.01112
More
README: all software under Apache 2.0. The GitHub licence API reports Apache-2.0 for the LICENSE file.
- Data licence
- One CC BY 4.0 notice covers all datasets. The terms at the original sources vary.1142
More
Blanket notice: 'All other materials' are under CC BY 4.0 (README, and a README.pdf in the data bucket). No per-dataset licences are listed.
The README also asks users to cite each contributed dataset. Whether the blanket notice replaces the original terms of contributed data is not stated. Not legal advice.
No licence column or files · The spreadsheet has no licence column. In the 101 dataset folders of gs://gresearch/robotics, the only licence file we found is DROID's (robotics/droid/1.0.0/CC-BY-4.0).1520
CC BY 4.0 at source · Stated at the source for BridgeData V2, DROID, RoboNet, VIMA (VimaBench README), RoboVQA and CLVR Jaco Play432044+3
CC BY-NC 4.0 at source (ManiSkill2 assets) · ManiSkill2 (30,000 simulated episodes in OXE) states its assets are under CC BY-NC 4.0 (README at tag v0.5.3).32
No data licence found at source · Language Table and SPOC repos state code licences (Apache-2.0) only; we found no separate data licence statement in their READMEs.3334
cc-by-4.0 (third-party mirror) · Hugging Face mirror jxu124/OpenX-Embodiment, 19,997 downloads (Hub 'downloads' field)48
- Asset licence
- unclear Inferred3215
More
Only the simulated datasets involve 3D assets. ManiSkill2 states its assets are CC BY-NC 4.0; OXE ships rendered observations, not the asset files, and no source says how the asset terms apply to renders. VIMA and SPOC asset terms were not checked.
- Access
- Open. The data is in a public storage bucket.112031
More
Public bucket gs://gresearch/robotics via TFDS or gsutil; no registration. The README's fallback download bucket (gs://gdm-robotics-open-x-embodiment) has data in only 35 of its 72 dataset folders; the other 37, including bridge and fractal20220817_data, hold only folder markers (our listing, 2026-10-10). See issues.i4.
- Commercial use
- Unclear Inferred111532+2
More
Code (Apache-2.0) and the blanket CC BY 4.0 notice both allow commercial use with attribution. But the pool re-hosts data from many labs; one component states non-commercial terms for its assets (ManiSkill2) and several state no data licence at all. So the whole pool is unclear; individual datasets may be clear. Not legal advice.
- Published at
- IEEE ICRA 2024, pages 6892-6903, DOI 10.1109/ICRA57147.2024.1061147737
More
Crossref record deposited by IEEE (published 2024-05-13). The arXiv page lists no venue.
Sources 48
- 1Open X-Embodiment paper, full text v9Paper · May 2025 · checked 10 Oct 2026
- 2Open X-Embodiment project siteOfficial site · Oct 2023 · checked 10 Oct 2026
- 3Evaluating Real-World Robot Manipulation Policies in Simulation (SimplerEnv), full text: Tables I, IV, V, XIIPaper · May 2024 · checked 10 Oct 2026
- 4Octo: An Open-Source Generalist Robot PolicyPaper · May 2024 · checked 10 Oct 2026
- 5OpenVLA: An Open-Source Vision-Language-Action Model (Sections 3.3, 5.1; Appendices A, B, C)Paper · Jun 2024 · checked 10 Oct 2026
- 6π0: A Vision-Language-Action Flow Model for General Robot ControlPaper · Oct 2024 · checked 10 Oct 2026
- 7π0.5: a Vision-Language-Action Model with Open-World GeneralizationPaper · Apr 2025 · checked 10 Oct 2026
- 8GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsPaper · Mar 2025 · checked 10 Oct 2026
- 9SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelPaper · Jan 2025 · checked 10 Oct 2026
- 10FAST: Efficient Action Tokenization for Vision-Language-Action Models (Appendix A, data mixture)Paper · Jan 2025 · checked 10 Oct 2026
- 11open_x_embodiment README (licence, download, RT-1-X checkpoint)Repository · Nov 2025 · checked 10 Oct 2026
- 12GitHub API: google-deepmind/open_x_embodiment (stars, forks, licence, open issues)Index · 10 Oct 2026 · checked 10 Oct 2026
- 13WorldGym: World Model as An Environment for Policy Evaluation (v3; Section 4.1, Appendix implementation details)Paper · May 2025 · checked 10 Oct 2026
- 14RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationPaper · Oct 2025 · checked 10 Oct 2026
- 15Open X-Embodiment Dataset Overview spreadsheet (read via CSV export)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 16VIMA: General Robot Manipulation with Multimodal Prompts (abstract: simulation benchmark, 600K+ expert trajectories)Paper · Oct 2022 · checked 10 Oct 2026
- 17SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World (abstract)Paper · Dec 2023 · checked 10 Oct 2026
- 18Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv abstract page, v1-v9 history)Paper · Oct 2023 · checked 10 Oct 2026
- 19Open X-Embodiment paper, PDF v9 (author list and affiliation footnote)Paper · May 2025 · checked 10 Oct 2026
- 20Google Cloud Storage listing of gs://gresearch/robotics (folders, timestamps, licence file search)Dataset page · 10 Oct 2026 · checked 10 Oct 2026
- 21Scaling up learning across many different robot types (blog)Official blog · 3 Oct 2023 · checked 10 Oct 2026
- 22Issue #30: discrepancy in the number of trajectories in the bridge dataset (author reply)Repository · Dec 2023 · checked 10 Oct 2026
- 23TensorFlow Datasets catalog: bridge (OXE copy, 25,460 train / 3,475 test)Dataset page · 10 Oct 2026 · checked 10 Oct 2026
- 24Octo dataset config (note: 'bridge_dataset' is not the official OXE copy)Repository · 2024 · checked 10 Oct 2026
- 25OpenVLA README (download Berkeley RLDS copy, rename to bridge_orig)Repository · Sep 2024 · checked 10 Oct 2026
- 26bridge_data_v2/0.0.1 dataset_info.json in gs://gresearch/robotics (60,063 episodes)Dataset page · 12 Aug 2024 · checked 10 Oct 2026
- 27Issue #11: Roboturk has no True values under is_last or terminate_episode (author confirmation)Repository · Oct 2023 · checked 10 Oct 2026
- 28Issue #85: Saytap dataset images are all zerosRepository · Jul 2024 · checked 10 Oct 2026
- 29open_x_embodiment issue tracker (incl. #45 failed fractal episodes, #95 missing berkeley_rpt views)Repository · 10 Oct 2026 · checked 10 Oct 2026
- 30Issue #104: Only 27 of 55 datasets available from the fallback bucketRepository · Jul 2025 · checked 10 Oct 2026
- 31Google Cloud Storage listing of gs://gdm-robotics-open-x-embodiment (README fallback bucket)Dataset page · 10 Oct 2026 · checked 10 Oct 2026
- 32ManiSkill README at tag v0.5.3 (ManiSkill2; licence section: assets CC BY-NC 4.0)Repository · 2023 · checked 10 Oct 2026
- 33Language Table repository (Apache-2.0; no separate data licence statement in README)Repository · Oct 2026 · checked 10 Oct 2026
- 34SPOC training repository (LICENSE: Apache 2.0)Repository · Nov 2024 · checked 10 Oct 2026
- 35Issue #24: Will there be RT-2-X weights? (author reply)Repository · Nov 2023 · checked 10 Oct 2026
- 36Issue #52: When will the code for RT-2 be available? (author reply)Repository · Mar 2024 · checked 10 Oct 2026
- 37Crossref record for the ICRA 2024 paperIndex · May 2024 · checked 10 Oct 2026
- 38bridge_data_msr dataset_info.json (Microsoft Research WidowX data in BridgeData V2 format)Dataset page · 11 Apr 2024 · checked 10 Oct 2026
- 39robo_ai_u_r5e dataset_info.json (UR5e data, Satakunta University of Applied Sciences)Dataset page · 27 Apr 2026 · checked 10 Oct 2026
- 40open_x_embodiment commit historyRepository · 5 Nov 2025 · checked 10 Oct 2026
- 41DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (v2; OXE scene count and OXE co-training baseline)Paper · Mar 2024 · checked 10 Oct 2026
- 42README.pdf in the data bucket (copyright and licence notice)Dataset page · Oct 2023 · checked 10 Oct 2026
- 43BridgeData V2 project site (licence statement)Official site · Aug 2023 · checked 10 Oct 2026
- 44RoboNet README (data under CC BY 4.0)Repository · Mar 2023 · checked 10 Oct 2026
- 45VimaBench README, licence table (dataset CC BY 4.0)Repository · Sep 2023 · checked 10 Oct 2026
- 46RoboVQA README, licence and disclaimerRepository · Dec 2023 · checked 10 Oct 2026
- 47GitHub API: clvrai/clvr_jaco_play_dataset (licence CC-BY-4.0)Index · 10 Oct 2026 · checked 10 Oct 2026
- 48Hugging Face: jxu124/OpenX-Embodiment (third-party mirror; licence label and downloads)Secondary · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
validity / sim_to_real: OXE paper v9 (no sim experiments). SimplerEnv (2405.05941) Tables I, IV, XII. WorldGym (2506.00613 v3). AutoEval (2503.24278), PolaRiS (2512.16881), REALM (2512.19562), 2606.10366, RoboWorld (2607.01060) and Ctrl-World (2510.10125): these concern Bridge or DROID and are recorded there. RobotArena ∞ (2510.23571): no real correlation measured. WorldEval (2505.19017) and dWorldEval (2604.22152): checked, they do not use OXE data. Leads from research/raw/sweeps/recent.json followed.
license_data (constituent datasets): OXE README and README.pdf in the bucket; spreadsheet columns; all 101 folders of gs://gresearch/robotics searched for licence files; READMEs or licence APIs of RoboNet, VimaBench, VIMA, RoboVQA, CLVR Jaco Play, Language Table, SPOC, ManiSkill (tags v0.4.2, v0.5.0, v0.5.3), TOTO, MimicPlay, RoboHive, Mobile ALOHA, DobbE, FurnitureBench, FMB. The last six give code licences only; not used as evidence of data terms.
objects, scenes: Paper v9 text and figures captions, project site, README, spreadsheet columns.
uncertainty_reported: OXE paper v9 HTML and PDF text: no 'standard error', 'confidence' or error-bar wording.
latest_update: GitHub commits, tags and releases; arXiv version history; bucket listings of gs://gresearch/robotics and gs://gdm-robotics-open-x-embodiment with object timestamps; spreadsheet.
Change history
- Created at full depth from primary sources, starting from the checked basic entry and research/raw/inventory/core-real.json. Re-checked every basic fact. Added: v1.0/v1.1 totals and the simulated share, unlisted bucket additions (2024, 2026), licence check of constituent datasets, bucket audit (only licence file is DROID's; fallback bucket half empty), RT-X result tables, Table I caption contradiction, undisclosed RT-2-X Bridge decoding workaround (per OpenVLA), validity list (SimplerEnv, offline MSE, WorldGym), adoption by 8 model reports. Corrected: blog says 33 labs (a fourth institution count); status rests on a 2026-04 bucket addition.
- Published as a full entry.