DROID
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
DROID is a dataset of 76k real robot demonstrations recorded in 564 scenes on one shared robot setup. It has no fixed test. Labs use the same setup to test policies (the models that control robots).123
What a score here does not tell you Inferred
- How one study's result compares with another study's result.There is no fixed test. Each study uses its own.
- Whether a low error on recorded data means success on a real robot.Action error (how far a policy's actions are from the recorded ones) correlated negatively with real results.
- How a policy would do on other kinds of robots.All of the data comes from one Franka setup.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| PolaRiS (simulated DROID scenes) Dec 2025 | Pearson correlation r = 0.90 and MMRV (a measure of how often two rankings disagree) = 0.03. The worst scene had r = 0.81.15 | The same 5 DROID-trained policies were scored in 6 reconstructed scenes and on real robots at UW and Princeton, with 20 real rollouts (test runs) per policy and scene. The study’s authors described the result as “strong correlation”. | The benchmark’s authors |
| PolaRiS compared with RoboArena Dec 2025 | Pearson r = 0.98 and MMRV = 0.0015 | The PolaRiS scores of 4 policies were compared with their average progress scores on RoboArena. The study’s authors described the result as “strong performance correlation”. | The benchmark’s authors |
| Offline action error (PolaRiS baseline) Dec 2025 | Pearson r = -0.55 on the training set and -0.53 on the validation set. MMRV = 0.40.15 | The same 5 policies were ranked by action error on the training set and on the validation set, and the rankings were compared with real performance. The study’s authors described the result as “poor metric”. | The benchmark’s authors |
| Ctrl-World (world model, own study) Oct 2025 | No coefficient was reported. The regression slopes were 0.87 for instruction following and 0.81 for success. By our calculation, Pearson r = 0.97 and 0.83 over 21 task-policy pairs.16 | DROID checkpoints of π0, π0-FAST and π0.5 were run on 7 tasks from the same start images, in a DROID-trained world model (a model that predicts what the cameras will see next) and on a real DROID setup. The study’s authors described the result as “closely correlated”. | The benchmark’s authors |
| Ctrl-World as tested by PolaRiS Dec 2025 | Pearson r = 0.53 and MMRV = 0.2215 | The same 5 policies were run in Ctrl-World across the 6 scenes of PolaRiS, scored by humans and compared with real results. The study’s authors described the result as “clear policy mis-rankings”. | The benchmark’s authors |
| REALM (Isaac Sim, own study) Dec 2025 | Overall Pearson r = 0.92, MMRV = 0.118 and p < 0.00117 | GR00T N1.5, π0 and π0-FAST ran 7 tasks under 5 perturbations in simulation and on a real DROID setup. There were about 800 paired rollouts, scored by task progress. The study’s authors described the result as “strong proxy”. | An independent group |
| Independent re-test of REALM, SIMPLER and VLA-Arena Jun 2026 | REALM had a mean Spearman of 0.700, a mean Pearson of 0.785 and a mean MMRV of 0.030. SIMPLER had 0.400, 0.402 and 0.128. VLA-Arena had 0.575, 0.725 and 0.060.18 | 5 policies (π0, π0-FAST, π0.5, GR00T N1.6 and GR00T N1.7) ran 9 matched tasks in three simulators and on real DROID hardware. There were 11,800 simulated and 1,115 real rollouts. | An independent group |
| RoboWorld (world model) compared with RoboArena Jul 2026 | Pearson r = 0.989 and Spearman = 0.97019 | 8 open policies were run in a DROID-trained world model from RoboArena start frames, with 4,186 rollouts. The results were compared with the RoboArena leaderboard. The study’s authors described the result as “align strongly”. | An independent group |
Benchmarks built on DROID: RoboArena, PolaRiS, REALM, Ctrl-World, RoboWorld, RobotArena ∞ DROIDSim.31517+3
Our assessment Opinion
DROID matters as a shared robot setup more than as a test. Always name the protocol behind a DROID number.
Reasoning
DROID matters for evaluation as a shared robot setup more than as a dataset. Because many labs own the same hardware, policies trained on it can be compared in the real world (RoboArena) and in calibrated simulations. A DROID-platform number is only meaningful with its protocol named.
Confidence: high
Many studies have checked proxies against real robots. The studies are small and they conflict.
Reasoning
DROID has the most paired evidence comparing simulation and real robots of any setup we have checked. The studies are small (3 to 8 policies) and mostly by overlapping teams, and they disagree when one group tests another's proxy. Treat any single proxy score as a rough indication.
Confidence: medium
Use version 1.0.1 together with the fix files released later.
Reasoning
Before training on DROID, use version 1.0.1 with the separate annotation, calibration and idle-filter files. Results from older pipelines are not directly comparable.
Confidence: medium
Known problems 8, 1 disputed
Some raw video is missing after face blurring
Some raw video episodes are missing. A 2024 note about this was never resolved.212223+1
Details
Since 2024-03-19 the docs have said 20% of raw episodes were missed when face-blurring and copying the raw data, 'to be fixed within a few days'. The note is still there. In 2025-03 a maintainer said a few episodes were likely missed in the extra face-blurring step for the non-RLDS data. A raw 1.0.1 folder dated 2024-03-18 exists, but the docs do not say whether it is complete. The RLDS version is complete per the docs.
Licences are missing for the code and later data folders
The control code has no licence. Copies of the data carry a different licence label.121425+2
Details
The hardware and control code repo has no licence file; a maintainer pointed users to the data licence in the bucket (2025-07). The CC BY 4.0 file sits only in the 1.0.0 folder. Hugging Face ports label the data Apache-2.0. A user asked in 2026-09 whether derived labels may be redistributed; no answer yet.
Proxy evaluations disagree with each otherDisputed
A proxy (a simulated or learned stand-in for real tests) scores well in its own paper. It scores worse when other groups test it.161519+2
Other view: Each proxy's authors report strong agreement in their own setup.151719
Details
World models: Ctrl-World's own data gives r 0.83 (our calculation), but PolaRiS measured it at r 0.53 with clear mis-rankings; RoboWorld reports r 0.989 on RoboArena. Simulators: REALM reported r 0.92 on its own; an independent group measured r 0.785 for REALM and 0.402 for a SIMPLER-style setup on DROID hardware. Each study used 3 to 8 policies.
PolaRiS, REALM and RoboWorld each report strong agreement within their own setups (r 0.90, 0.92, 0.989).
Offline action error is a poor predictor of real performance
Lower action error did not mean better real performance.159
Details
In PolaRiS, action error on training and validation data correlated negatively with real performance of 5 DROID policies (r -0.55 and -0.53). NVIDIA's GR00T DROID example still reports open-loop action error as its expected performance figure.
Counts differ between sources and versions
The counts of tasks, trajectories and episodes differ between sources.28129+3
Details
Tasks: 86 (paper body, site, RSS) vs 84 (arXiv abstract). Trajectories: 76k (arXiv, site) vs 65k (RSS abstract). Episodes: 92,233 (RLDS 1.0.0) vs 95,658 (1.0.1); Ctrl-World cites 95,599. Scene types: 10 (Section IV) vs 9 (Fig. 10). Collecting groups: 13 institutions vs 18 labs.
Some models could not fit DROID in large training mixtures
Some models dropped DROID from their training mix because they could not fit it.674+1
Details
OpenVLA trained with DROID at 10% weight but removed it for the last third of training because action-token accuracy stayed low; SpatialVLA did the same. The FAST paper says OpenVLA struggled to fit the higher-frequency DROID data and that FAST enabled the first strong generalist DROID policy. RobotArena ∞ states DROID is often left out of pretraining because of higher noise, citing a third paper.
Data fixes were released after the dataset
Camera calibrations, language labels and idle frames (pauses where the robot does not move) needed fixes after the release.32332
Details
The official annotations repo says the original release 'included noisy camera calibrations' (better ones for about 36k episodes since April 2025), that the released RLDS contains only a subset of the language labels (full labels since December 2024), and that many episodes contain long pauses, mostly at the start, which make policies output idle actions. It recommends filtering them with a list valid only for version 1.0.1 (added August-September 2025). Physical Intelligence's openpi guide says idle filtering significantly improves policy performance.
There is no standard test on DROID hardware
DROID results come from test protocols that are not compatible with each other.143
Details
Results come from different protocols: the paper's co-training A/B tests (6 tasks), FAST's 44-trial suite, RoboArena's crowd-sourced pairwise ratings, and simulated proxies scored by task progress. RoboArena found that a conventional single-lab suite (FAST's) ranked policies less accurately than distributed evaluation (r 0.692 vs 0.838 with an exhaustive ranking, under simulated shifts).
Details
About
- What it is
- Dataset12
More
Released as a dataset with training code, checkpoints and a hardware guide. No fixed test or scoring rule.
- Built by
- Stanford University, UC Berkeley, Toyota Research Institute, 13 collecting institutions128
More
Authors · 101 authors; first authors Alexander Khazatsky and Karl Pertsch. 16 numbered affiliations including Stanford, UC Berkeley, Toyota Research Institute, CMU, UT Austin, Princeton, University of Washington, Google DeepMind and KAIST.281
Count of collecting groups differs · '13 institutions' and '18 robots' (Sections I and III) vs '18 research labs' (introduction). Raw data folders are split by 13 lab names.124
- Released
- March 2024, at RSS 2024283029
More
arXiv v1 2024-03-19; RLDS 1.0.0 files in the bucket dated 2024-03-15. Published at RSS 2024.
- Version
- RLDS 1.0.1303124+2
More
RLDS 1.0.0 (92,233 episodes) and 1.0.1 (95,658 episodes); raw 1.0.0 and 1.0.1 (face-blurred stereo video); droid_100 sample (100 episodes). Extra annotations ship separately on Hugging Face.
1.0.0 vs 1.0.1 language labels · openpi: 1.0.1 has the complete language annotations (about 75k episodes), 1.0.0 only 30k. The official annotations repo says the released RLDS dataset contains only a subset of labels.3332
OXE copy · OXE v1.1 lists DROID with 92,233 episodes (1,670 GB), the 1.0.0 count.34
- Last update
- September 2025. A data filter was updated.323531
More
2025-09-01: updated idle-frame filter in the official annotations repo. 2025-09-15: last repo commit (docs and site config). RLDS 1.0.1 objects in the bucket carry 2025-07-15 timestamps.
Site news: December 2024 language annotations; April 2025 improved calibrations for 36k episodes; arXiv v2 2025-04-22. A Hugging Face port named droid_1.0.1 existed from 2025-03-17, so version 1.0.1 predates its bucket timestamps.
- Status
- No official updates since September 2025. Use is very active. Inferred353212+1
More
No official update since September 2025. Maintainers answered issues in 2025. Use as training data and test platform is very active.
54 open issues and PRs; users still file data questions in 2026 (e.g. #80 on redistributing derived labels, #73 stereo availability). Promised scene-type metadata (2024-04) has not appeared in a dataset version.
Setup
- Runs in
- Real robots1
- Robot
- One arm1
- Robot model
- Franka Panda136
More
Franka Panda 7-DoF arm with Robotiq 2F-85 gripper on a movable standing desk; two Zed 2 stereo cameras and a wrist Zed Mini; Oculus Quest 2 teleoperation; 15 Hz
The hardware list puts the total cost at about $46,000 (updated around September 2025; earlier $20,000), mostly the arm.
- Setting
- Whole home, Office or lab, Kitchen, Mixed1
More
The paper says 10 scene types (Section IV) and 9 (Fig. 10 caption).
- Tasks
- 86 tasks, counted as verbs12928
More
86 tasks, counted as unique verbs in the instructions (paper body, site, RSS abstract). The arXiv abstract says 84.
CONFLICT between the arXiv abstract and the body. A verb count is not comparable with task counts of fixed benchmarks.
- Training data
- 76k demonstrations, 350 hours129
More
76k successful demonstrations (350 hours) by 50 collectors over 12 months. About 16k trajectories marked 'not successful' are also released. The RSS abstract says 65k.
CONFLICT: 76k (arXiv, site) vs 65k (RSS 2024 abstract). RLDS episode counts include failures.
RLDS episode counts · 1.0.0: 92,233 episodes (1,834,749,018,029 bytes). 1.0.1: 95,658 episodes (1,865,993,126,270 bytes).3031
Up to 3 crowdsourced instructions per episode · Labelled after collection on the tasq.ai platform; 3 labels for 95% of the 75k successful episodes since December 2024.12
- Size
- 1.7 TB (RLDS)21
More
RLDS 1.7 TB; raw stereo HD video 8.7 TB; raw non-stereo 5.6 TB (docs)
A user reports the non-stereo raw subset is 3.9 TB (issue #60, open).
Scoring and access
- Scored by
- Success rate1
More
Later DROID-platform evaluations use progress scores (PolaRiS, REALM) or pairwise preferences (RoboArena).
- Score
- Success rate. Later tests use other scores.1
More
Paper: diffusion policies trained on small in-domain data, with or without 50/50 co-training on DROID or OXE; 6 tasks at 4 locations, in-distribution and out-of-distribution. DROID beat the next best by 22 points (in distribution) and 17 points (OOD).
Subset used · The paper's policies used the first 40K successful trajectories that had language labels at the time.1
FAST 'DROID evaluation' · First zero-shot test of DROID policies in unseen scenes (2025-01): 44 trials per policy over 17 tasks in its Table II (its text says 16).4
RoboArena · Pairwise blind A/B tests of DROID policies by evaluators at many sites; ratings from a task-aware Bradley-Terry model. Has its own record.3
Offline action error · NVIDIA's GR00T DROID example reports open-loop action error on DROID episodes (average MSE about 0.0149) as expected performance.9
- Trials
- 10 per task setting143+1
More
Paper: 10 rollouts per task setting and method. FAST: 44 per policy. RoboArena paper: 4,284 episodes. PolaRiS: 20 real rollouts per policy and environment.
- Who runs it
- Each team tests its own model Inferred13
More
Papers run their own DROID-platform trials. RoboArena, a separate project, collects crowd-sourced evaluations of submitted policies.
- Error bars
- Sometimes reported Inferred1411
More
The paper reports standard errors; FAST reports 95% intervals; RoboArena shows a standard deviation per rating. Many later reports give single numbers.
- Leaderboard
- None. Scores are only in papers. Inferred211
More
No leaderboard on the DROID site. RoboArena keeps a live DROID-platform leaderboard (24 policies on 2026-10-11 UTC).
- Code licence
- None for the control code. MIT for the policy code.121314
More
The hardware and control code repo has no LICENSE file (GitHub licence API: null). The policy-learning repo is MIT (copyright 2021 Stanford Vision and Learning Lab, a robomimic fork).
In issue #62 (2025-07) a user asked for a licence; a maintainer pointed to the data licence in the bucket. No licence was added to the repo.
- Data licence
- CC BY 4.012514
More
CC BY 4.0 (paper; licence file in the bucket next to RLDS 1.0.0)
There is no licence file in the 1.0.1 folder. Hugging Face ports carry a different label (items); they do not change the original terms.
apache-2.0 · lerobot/droid_1.0.1 (Hugging Face; 22,682 downloads) and cadene/droid_1.0.1 (163,135 downloads), Hub 'downloads' field2637
- Asset licence
- not applicable Inferred21
More
Real-robot recordings only. People's faces in the raw video were blurred before release (docs; internal name r2d2_faceblur).
- Access
- Open, in a public storage bucket2124
More
Public bucket gs://gresearch/robotics/droid via TFDS or gsutil; no registration.
- Commercial use
- Unclear Inferred251213
More
The data (CC BY 4.0) and policy code (MIT) allow commercial use with attribution. The hardware and control software repo has no licence, so reuse rights for it are not granted. The dataset alone reads as allowed. Not legal advice.
- Published at
- Robotics: Science and Systems XX, paper 120, July 2024, DOI 10.15607/RSS.2024.XX.12029
Sources 37
- 1DROID paper, full text v2Paper · Apr 2025 · checked 10 Oct 2026
- 2DROID project site (updates, Hugging Face link)Official site · Apr 2025 · checked 10 Oct 2026
- 3RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (v2)Paper · Jun 2025 · checked 10 Oct 2026
- 4FAST: Efficient Action Tokenization for Vision-Language-Action Models (DROID evaluation, Table II)Paper · Jan 2025 · checked 10 Oct 2026
- 5openpi README (DROID checkpoints)Repository · 2026 · checked 10 Oct 2026
- 6OpenVLA (Section 3.3, Appendix A; Franka-DROID tests)Paper · Jun 2024 · checked 10 Oct 2026
- 7SpatialVLA (DROID removed for final third of pretraining)Paper · Jan 2025 · checked 10 Oct 2026
- 8GR00T N1 (DROID among OXE subsets, 23.1M frames)Paper · Mar 2025 · checked 10 Oct 2026
- 9Isaac-GR00T DROID example (GR00T-N1.7-DROID; open-loop MSE)Repository · 2026 · checked 10 Oct 2026
- 10V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningPaper · Jun 2025 · checked 10 Oct 2026
- 11RoboArena leaderboard API (all policies)Leaderboard · 11 Oct 2026 · checked 10 Oct 2026
- 12GitHub API: droid-dataset/droid (licence null, stars, issues)Index · 10 Oct 2026 · checked 10 Oct 2026
- 13droid_policy_learning LICENSE (MIT)Repository · 2024 · checked 10 Oct 2026
- 14droid issue #62: License (maintainer reply)Repository · Jul 2025 · checked 10 Oct 2026
- 15PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Figures 6-8, 13)Paper · Dec 2025 · checked 10 Oct 2026
- 16Ctrl-World: A Controllable Generative World Model for Robot Manipulation (v3; Table 3, Figure 7)Paper · Oct 2025 · checked 10 Oct 2026
- 17REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation (PDF v1, Fig. 6)Paper · Dec 2025 · checked 10 Oct 2026
- 18A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation (Table 2)Paper · Jun 2026 · checked 10 Oct 2026
- 19RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation (v4)Paper · Jul 2026 · checked 10 Oct 2026
- 20RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationPaper · Oct 2025 · checked 10 Oct 2026
- 21DROID docs: The DROID Dataset (download sizes, schema, raw-data note)Official site · Mar 2025 · checked 10 Oct 2026
- 22History of docs/the-droid-dataset.md (raw-data note added 2024-03-19)Repository · 11 Mar 2025 · checked 10 Oct 2026
- 23droid issue #47: No corresponding raw data for many RLDS episodes (maintainer reply)Repository · Mar 2025 · checked 10 Oct 2026
- 24Google Cloud Storage listing: gs://gresearch/robotics/droid, droid_raw, droid_100Dataset page · 10 Oct 2026 · checked 10 Oct 2026
- 25CC-BY-4.0 licence file in the DROID 1.0.0 bucket folderDataset page · 15 Mar 2024 · checked 10 Oct 2026
- 26Hugging Face: lerobot/droid_1.0.1 (licence label, downloads)Repository · Jun 2026 · checked 10 Oct 2026
- 27droid issue #80: permission to redistribute derived labelsRepository · Sep 2026 · checked 10 Oct 2026
- 28DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (arXiv abstract page)Paper · Mar 2024 · checked 10 Oct 2026
- 29DROID, RSS 2024 proceedings page (paper 120)Paper · Jul 2024 · checked 10 Oct 2026
- 30DROID RLDS 1.0.0 dataset_info.jsonDataset page · 15 Mar 2024 · checked 10 Oct 2026
- 31DROID RLDS 1.0.1 dataset_info.jsonDataset page · 15 Jul 2025 · checked 10 Oct 2026
- 32KarlP/droid on Hugging Face: official annotations, calibrations, idle filter (README and commits)Repository · 1 Sep 2025 · checked 10 Oct 2026
- 33openpi DROID training guide (version 1.0.1, idle filtering)Repository · Sep 2025 · checked 10 Oct 2026
- 34Open X-Embodiment dataset spreadsheet (DROID row)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 35droid repo commit historyRepository · 15 Sep 2025 · checked 10 Oct 2026
- 36DROID docs: hardware shopping list (approximate total cost)Repository · Sep 2025 · checked 10 Oct 2026
- 37Hugging Face: cadene/droid_1.0.1 (licence label, downloads)Repository · Mar 2025 · checked 10 Oct 2026
Where we searched for missing information
validity / sim_to_real: DROID paper v2 (no sim); PolaRiS v2 (Figures 6-8, 13); Ctrl-World v3 (Table 3, PDF Figure 7); REALM v1 (PDF Fig. 6); 2606.10366 (Table 2); RoboWorld v4; RobotArena ∞ (no real correlation); RoboArena v2 (real vs oracle, not sim); 2508.11117 (position paper, no measurement); WorldEval and dWorldEval (do not use DROID).
objects: Paper text and figure captions, site, docs: no total.
license_code, licence files: droid repo root and GitHub licence API; droid_policy_learning LICENSE; bucket folders droid/1.0.0, droid/1.0.1, droid_raw; issue #62.
top_score: No headline score for the dataset. RoboArena ratings belong to the RoboArena record.
Change history
- Created at full depth from primary sources, starting from the checked basic entry and research/raw/inventory/core-real.json. Added eight paired validity comparisons, post-release data fixes, raw-data gap, licence gaps, hardware cost, RoboArena use and adoption. Corrected the basic entry: latest update is 2025-09 (annotation repo), not 2025-07; the version-label claim is now verified at the official annotation repo; the paper names 10 and 9 scene types in different places.
- Published as a full entry.