RoboArena
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
RoboArena is a service that tests robot policies (the robots' control models) on real DROID robot arms. Volunteers run blind A/B tests, in which two unnamed policies try the same task, and their preferences are combined into a rating.123
What a score here does not tell you Inferred
- How a policy performs on other robots.RoboArena tests only the DROID Franka setup.
- How often a policy succeeds on a fixed set of tasks.Scores are relative ratings from tests where evaluators choose the tasks.
- How good the rarely tested policies are.Of the 24 policies, 15 have fewer than 100 evaluations.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| Oracle ranking check (paper Figure 6) Jun 2025 | Pearson r = 0.98 and MMRV = 1.8% for the task-aware model. A conventional 17-task lab protocol gave r = 0.69 and MMRV = 13%. MMRV measures how often two rankings disagree.212 | Rankings built from 612 A/B tests of 7 DROID policies were compared with an 'oracle' ranking. After each A/B test, the evaluator also ran the other 5 policies on the same task. The oracle ranks policies by their average progress score over 4,284 rollouts. The study’s authors described the result as “more accurately rank”. | The benchmark’s authors |
| Simulated drift re-analysis (paper Table 2, arXiv v2) Nov 2025 | Pearson r = 0.838 and MMRV = 0.058. The conventional protocol gave r = 0.692 and MMRV = 0.141.2 | The same data were ranked again after simulating a shift over time from easy to hard tasks and from weaker to stronger policies. The new rankings were compared with the oracle. The study’s authors described the result as “significantly higher correlation”. | The benchmark’s authors |
| Leave-one-organisation rerun (live audit tool) Jun 2026 | Without the largest label (57.6% of evaluations), ranks move by at most 2 places. Without 'nvidia', DreamZero falls below 100 evaluations.131415 | The organisers' site recomputes the official ranking with all evaluations from one evaluator organisation removed. We read the reruns for the two most influential labels on 2026-10-10. | The benchmark’s authors |
Our assessment Opinion
A rating shows how often evaluators preferred a policy in this pool, on this robot.
Reasoning
A RoboArena rating says how often evaluators preferred a policy over the others in the pool, on DROID Franka arms, in tasks the evaluators chose. It is a relative number that moves when the pool or the evaluators change, and it does not translate into a success rate or carry over to other robots.
Confidence: high
Check evaluation counts and evaluator shares before trusting the top places.
Reasoning
The top of the board rests on thin evidence. The official leader's place depends on evaluations carrying its developer's label, and the all-policies leader has 25 evaluations from essentially one label. Read the evaluation counts and the audit metrics next to every rating.
Confidence: high
The accuracy check uses only RoboArena's own data.
Reasoning
The paper's accuracy check compares RoboArena rankings with an oracle ranking (a reference built from the same evaluators, tasks and progress scores), over 7 related policies. It shows that the pairwise method recovers that oracle efficiently. It is not a comparison with an independent measure of real-world usefulness.
Confidence: medium
Open data made the manipulation visible. A few evaluators still supply a large share of the tests, which remains a risk.
Reasoning
Public evaluation data let the organisers and outsiders spot the spring 2026 manipulation, and the rollback can be traced in archived data. The fixes are rules and thresholds added afterwards. A large share of evaluations still comes from a few evaluators, so the board stays vulnerable to them.
Confidence: medium
Known problems 6
The rollback removed 847 evaluations and changed the top of the board
847 evaluations from April to June 2026 were removed. Spirit v1.6, which led the board before, now has 25 counted tests. Inferred161718+2
Details
Comparing the official data archived on 2026-06-03 with today's data: 847 of the 4,614 public evaluations then listed are gone, all dated 2026-04-02 to 2026-06-03. They came from six evaluator-organisation labels (333, 276, 142, 72, 22 and 2 evaluations); of the 954 evaluations listed for that period, 107 remain. Spirit v1.6 led with 1926 ± 37.3 from 310 evaluations; 223 of its 309 listed evaluations came from two labels that first appeared in April and May 2026 and recorded 212 wins, 1 loss and 10 ties for it, while all other evaluators recorded 33 wins, 37 losses and 16 ties. After the rollback it has 25 counted evaluations and appears only in the all-policies view. WALL-OSS (fourth on 2026-06-03, 48 evaluations) no longer appears. MolmoAct2-DROID went from 63 evaluations and 1066 to 5 evaluations and 1579. Press reports say Spirit AI had announced first place on 2026-06-03.
Most evaluations come from one evaluator label
One evaluator label supplied 57.6% of all counted evaluations and 92.9% since mid-June 2026. Inferred211813
Details
One evaluator-organisation label, 'frodobots', contributed 2,301 of 3,994 counted evaluations (57.6%) between 2025-08-02 and 2026-09-12. Its share was 659 of 987 in 2026 (66.8%) and 195 of 210 after 2026-06-13 (92.9%). For every official policy except DreamZero, frodobots supplied the largest share of its evaluations (53.9% to 67.5%). The official leave-one-organisation rerun without this label keeps the top two policies but moves others by up to two places.
The official leader was evaluated mostly under its developer's label
58% of the official leader's evaluations carry its developer's organisation label. Inferred182214+2
Details
DreamZero, an NVIDIA model, has 190 counted evaluations; 111 of them (58.4%) carry the evaluator label 'nvidia' and date from 2026-01-28 to 2026-02-26. These predate the 2026-04-02 start of the retroactive no-stake rule and remain counted. The site's own leave-one-organisation rerun without 'nvidia' leaves DreamZero with 79 evaluations, below the 100-evaluation threshold, so pi0.5-DROID would lead the official view. DreamZero has had no new evaluations since 2026-03-30 and its server was offline in the 30 days to 2026-10-11.
Ratings shift with the pool, and one weekly snapshot was broken
Scores move when other policies are added, and one weekly snapshot shifted every score by about 240 points. Inferred24
Details
Ratings are relative. In the published weekly snapshots, policies with no new evaluations still change score; for example Spirit v1.6 ranged from 1773 to 1792 between June and October 2026. On the 2026-08-24 snapshot every policy rose by 244 to 278 points and fell back by about 233 a week later; that snapshot gave a newly added policy with 2 evaluations a score of -3884 and a standard deviation of 39,859.4, and listed a standard deviation of about 1733 for every other policy.
Evaluations were manipulated in spring 2026
The organisers found manipulation, rolled back evaluations from April 2026 and changed the rules.232515+1
Details
On 2026-06-13 the RoboArena team said it had observed evidence of benchmark manipulation since April. Its notice says one sign was unusually low completion rates for requested evaluation assignments, and that organisations completing less than 20% of their requested evaluations were flagged. The team excluded evaluations from organisations with suspicious patterns and now allows only third-party evaluators with no stake in submitted policies. Both changes were applied retroactively to evaluations from 2026-04-02; the team says it found no evidence requiring exclusion before that date. It also added a 100-evaluation threshold for the official view, audit metrics and easier data downloads. The notice was removed from the site on 2026-07-17.
Limitations stated by the authors
The authors say RoboArena covers only the DROID platform and cannot easily run controlled tests that change one factor at a time. They did not study how robust it is to evaluators who act in bad faith.2
Details
The paper lists: evaluation on the DROID platform only; difficulty running controlled experiments that vary one factor at a time; robustness to intentionally adversarial evaluators not studied; possible over-optimisation (Goodhart's law); and possible latency from remote inference, found negligible for static tasks. It also reports that preference and progress feedback disagreed in 11% of A/B evaluations.
Details
About
- What it is
- Evaluation service Inferred13
More
Classified by the Atlas. Submitters host a policy server; RoboArena's evaluator network runs it on real robots. RobotArena (with the infinity sign) is a different project.
- Built by
- UC Berkeley, Stanford University, University of Washington, University of Montreal, NVIDIA, University of Pennsylvania, UT Austin, Yonsei University227
More
arXiv v2 lists 32 authors; the PMLR version lists 26. Eight numbered affiliations in arXiv v2.
UC Berkeley · Lead institution; two of three corresponding authors (Atreya, Pertsch).2
Stanford University · Third corresponding author (Tony Lee).2
University of Washington, University of Montreal, University of Pennsylvania, UT Austin, Yonsei University · Evaluator sites and co-authors in the paper.2
NVIDIA · Co-authors who built the simulated evaluation environments (Appendix A).2
- Released
- June 2025, at CoRL 202512721
More
arXiv v1 on 2025-06-22. Published at CoRL 2025 (PMLR volume 305). The first public A/B evaluation is dated 2025-04-15.
Code repository created 2025-06-14; website repository created 2025-06-17.
- Version
- No version numbers. The leaderboard is live, and there are 3 data dumps.282919
More
No versions, tags or releases. The leaderboard is recomputed continuously and three data dumps exist (2025-08-05, 2026-02-03, 2026-07-17).
Paper v1 (2025-06-22) and v2 (2025-11-29) · v2 adds Table 2 (simulated task and policy drift) and two authors. Figure 6 is identical in both.1212
Official and All views (since 2026-06-10) · The official view lists only policies with at least 100 A/B evaluations. The all-policies view adds entries with fewer, marked 'low sample'.315
Integrity update (2026-06-11 to 2026-06-13) · New evaluator rule, evaluations from 2026-04-02 rolled back, audit metrics added. See issues.i1.231525
- Last update
- The leaderboard was recomputed in October 2026. The newest evaluation is from 12 September 2026.192130+1
More
Leaderboard recomputed 2026-10-11 03:30 UTC (evening of 2026-10-10 in the US). Newest public A/B evaluation 2026-09-12. Latest data dump 2026-07-17. Code last changed 2026-04-28.
- Status
- Active. The site says it will run through December 2026. Inferred183
More
Evaluations every month from April to September 2026, at a lower rate than in 2025. The site says the benchmark runs live through December 2026, with possible extensions.
Counted evaluations per month in 2026 (our count from the public list): Jan 234, Feb 237, Mar 171, Apr 74, May 43, Jun 85, Jul 68, Aug 60, Sep 15 (to 2026-09-12). None between 2026-09-12 and 2026-10-10. In 2025 there were 3,008. Four of 24 policy servers were online in the 30 days to 2026-10-11.
Setup
- Runs in
- Real robots2
More
Closed-loop runs on physical DROID robots.
- Robot
- One arm2
More
One Franka arm on the DROID platform. The authors list evaluation on DROID only as a limitation.
- Robot model
- DROID platform: Franka Panda 7-DoF arm, Robotiq 2F-85 gripper, ZED-mini wrist camera, one or more ZED 2 external cameras, on a height-adjustable mobile table.2
- Setting
- Mixed Inferred2
More
Evaluators use their own labs, offices and kitchens. The paper reports dozens of scenes and shows 32 sample environments (Figure 11).
- Tasks
- Open-ended, with 2,797 distinct instructions Inferred218
More
No fixed task list. The paper reports hundreds of instructions; the public evaluations use 2,797 distinct instruction strings.
Distinct strings counted by us on 2026-10-10; after lower-casing and removing punctuation, 2,761. Most frequent: 'open the book' (40), 'close the book' (35).
- Training data
- None. Submitters use DROID data.310
More
RoboArena supplies no training data. It points submitters to the open DROID dataset and the openpi DROID training code.
- Size
- 3,994 A/B evaluations of 24 policies21
More
Paper (2025) · 7 policies, 612 A/B comparisons, 4,284 evaluation rollouts including the oracle runs, 7 universities.2
Live service (2026-10-10) · 3,994 counted A/B evaluations (575 ties, 14.4%), 24 policies, 120 policy pairs, 60 evaluator accounts, 28 evaluator-organisation labels. First evaluation 2025-04-15, last 2026-09-12.21
Data dump 2026-07-17 · 3,883 evaluation sessions, 10,783 policy episodes, 27,148 videos, 21.7 GB (21,676,174,280 bytes).30
Removed in June 2026 · 847 evaluations dated 2026-04-02 to 2026-06-03 were taken out of the public list and the ranking (issues.i2).1718
- Changes at test
- Open-ended. Evaluators choose the scenes and tasks. Inferred218
More
Evaluators pick any scene, objects, camera view and instruction. Variation is wide but uncontrolled and not labelled.
Paper Figure 11: lighting, camera viewpoints, tablecloths and objects were often changed. The 3,995 public evaluations use 2,797 distinct instruction strings (our count). RoboArena does not record whether a task or object resembles the policy's training data.
Scoring and access
- Scored by
- Human preference, Progress score218
More
For each A/B test the evaluator gives a preference (A, B or tie in the live data), a 0 to 100 progress score per policy, and a written reason. The paper found preference and progress disagree in 11% of A/B evaluations. The live rating is described as a Bradley-Terry-Davidson fit to preferences; we did not find whether progress scores enter it.
- Score
- A rating from pairwise comparisons, with a standard deviation. The official view lists only policies with at least 100 tests.3192
More
Each policy gets a rating on an Elo-like scale (786 to 1788 on 2026-10-10) with a standard deviation, fitted to all counted A/B preferences, ties included. The official view lists only policies with at least 100 A/B evaluations.
The paper's preferred model adds latent task-difficulty buckets fitted by an EM algorithm (Section 3.2, Appendix B). The site calls the live method a 'Bradley-Terry Davidson ranking path' and labels its chart 'Elo score'. Ratings are relative: other policies' results move every value (issues.i5). Comparisons involving the base pi0 or pi0-FAST models are shown but not counted.
- Trials
- Each policy runs once in each A/B test23
More
Each A/B evaluation runs each of the two policies once, back to back, from the same start. The official view needs at least 100 A/B evaluations per policy. In the paper's oracle sessions, evaluators ran all 7 policies on each task.
- Who runs it
- The organisers run the tests Inferred2323
More
Step by step: (1) A submitter hosts the policy on a remote server and submits it through a form; the team checks the input and output format and runs it with a trained evaluator in a test setup before adding it to the pool. (2) An evaluator's client asks the central server for a pair and receives the two servers' IP addresses, not the policy names. (3) The evaluator arranges a scene, types an instruction and runs policy A and then policy B from the same start until success or a timeout. (4) The evaluator records a preference, two progress scores and a reason; videos and actions are uploaded. (5) The central server fits the rating. Evaluators earn evaluation credits; submitters get a weekly budget. Since June 2026 only evaluators with no stake in submitted policies may volunteer. Taxonomy fit is loose: the organisers coordinate but do not run the robots.
- Error bars
- Usually reported191231
More
The leaderboard shows a standard deviation for every rating, and the paper's figures carry error bars. PhAIL (2026-05) says neither RoboArena nor RoboChallenge reports confidence intervals or paired tests on the rankings.
- Leaderboard
- Official leaderboard193
More
Public leaderboard at robo-arena.github.io/leaderboard, with an A/B evaluation viewer and per-organisation audit metrics.
- Code licence
- MIT91011
More
LICENSE: 'Copyright (c) 2025 robo-arena'. It covers the public repository, which holds the policy-server interface, client example and a test script. The evaluator client, central server and ranking code are not in it (file tree checked 2026-10-10). The website repository and the linked DROID simulation repository (arhanjain/sim-evals) have no licence file.
- Data licence
- MIT2930
More
All three Hugging Face data dumps carry license: mit. The paper text is CC BY 4.0 on arXiv.
- Asset licence
- not applicable Inferred2
More
Tests use physical robots and objects; no 3D assets are distributed.
- Access
- Policies are submitted through a form. The data is open.32330
More
Policies are submitted through an online form and checked before they enter the pool; evaluators also sign up through a form. Results, videos and data dumps are open to all.
- Commercial use
- Allowed Inferred930
More
Code and data dumps are MIT-licensed. Not legal advice. The dumps contain videos recorded in evaluators' labs.
- Published at
- CoRL 2025 (PMLR 305:336-364)27
More
arXiv comments give no venue; the PMLR page confirms.
Sources 31
- 1RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv abstract page, v1 2025-06-22, v2 2025-11-29)Paper · Jun 2025 · checked 10 Oct 2026
- 2RoboArena paper, full text v2 (Sections 3-7, Appendices B-D)Paper · Nov 2025 · checked 10 Oct 2026
- 3RoboArena website and its page code (about text, leaderboard views, alias table, audit metrics)Official site · 17 Jul 2026 · checked 10 Oct 2026
- 4PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Figure 8, correlation with RoboArena)Paper · Dec 2025 · checked 10 Oct 2026
- 5RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies (v4; Fig. 10, RoboArena Elo comparison)Paper · Apr 2026 · checked 10 Oct 2026
- 6RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation (v4; RoboArena leaderboard as ground truth)Paper · Jul 2026 · checked 10 Oct 2026
- 7openpi DROID example README (RoboArena baseline checkpoints; pi0.5-DROID named strongest based on RoboArena)Repository · 2026 · checked 10 Oct 2026
- 8Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models (NVIDIA Technical Blog)Official blog · 15 Jun 2026 · checked 10 Oct 2026
- 9robo-arena/roboarena LICENSE fileRepository · Jun 2025 · checked 10 Oct 2026
- 10robo-arena/roboarena README and file tree (policy server interface only)Repository · 28 Apr 2026 · checked 10 Oct 2026
- 11GitHub API: arhanjain/sim-evals (DROID simulated evaluation linked by RoboArena; no licence)Index · 10 Oct 2026 · checked 10 Oct 2026
- 12RoboArena paper Figure 6: ranking accuracy versus the oracle (Pearson r and MMRV bar labels; identical in v1 and v2)Paper · Jun 2025 · checked 10 Oct 2026
- 13RoboArena official leave-one-organisation rerun: label 'frodobots' removedLeaderboard · Oct 2026 · checked 10 Oct 2026
- 14RoboArena official leave-one-organisation rerun: label 'nvidia' removedLeaderboard · Oct 2026 · checked 10 Oct 2026
- 15RoboArena website repository commit history (notice added 2026-06-11; 100-evaluation official view 2026-06-10; notice hidden 2026-07-17)Repository · 17 Jul 2026 · checked 10 Oct 2026
- 16RoboArena leaderboard data as archived on 2026-06-03 (last_updated 2026-06-03 11:49 UTC), before the rollbackLeaderboard · 3 Jun 2026 · checked 10 Oct 2026
- 17RoboArena public A/B evaluation list as archived on 2026-06-03 (4,614 sessions), before the rollbackLeaderboard · 3 Jun 2026 · checked 10 Oct 2026
- 18RoboArena public A/B evaluation list (3,995 sessions with instruction, preference, progress scores and evaluator organisation label)Leaderboard · Oct 2026 · checked 10 Oct 2026
- 19RoboArena leaderboard data (JSON loaded by robo-arena.github.io/leaderboard; last_updated 2026-10-11 03:30 UTC)Leaderboard · Oct 2026 · checked 10 Oct 2026
- 20Huxiu: report on RoboArena manipulation and Spirit v1.6 (republished from WeChat account Dianchang)Secondary · 24 Jun 2026 · checked 10 Oct 2026
- 21RoboArena public transparency statistics (counted A/B evaluations by policy, pair and evaluator organisation)Leaderboard · Oct 2026 · checked 10 Oct 2026
- 22World Action Models are Zero-shot Policies (DreamZero), arXiv 2602.15922 v1Paper · Feb 2026 · checked 10 Oct 2026
- 23RoboArena benchmark integrity notice, website source at commit a6602e6 ('Publish post-maintenance site')Official site · 13 Jun 2026 · checked 10 Oct 2026
- 24RoboArena weekly leaderboard snapshots (18 snapshots, 2026-06-10 to 2026-10-05)Leaderboard · Oct 2026 · checked 10 Oct 2026
- 25Post by Pranav Atreya (RoboArena lead author) on benchmark hacking and rolled-back evaluationsOfficial blog · 13 Jun 2026 · checked 10 Oct 2026
- 26What Do Robotics Leaderboards Tell Us About The State of Robot Learning? (It Can Think! newsletter)Secondary · 13 Jun 2026 · checked 10 Oct 2026
- 27RoboArena, Proceedings of the 9th Conference on Robot Learning, PMLR 305:336-364Paper · Sep 2025 · checked 10 Oct 2026
- 28GitHub API: robo-arena/roboarena (stars, forks, created, pushed; no releases or tags)Index · 10 Oct 2026 · checked 10 Oct 2026
- 29Hugging Face API: datasets by RoboArena (three data dumps, licence tags)Repository · 17 Jul 2026 · checked 10 Oct 2026
- 30RoboArena/DataDump_07-17-2026 dataset card and Hub record (size, downloads)Repository · 17 Jul 2026 · checked 10 Oct 2026
- 31PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology (Related Work)Paper · May 2026 · checked 10 Oct 2026
Where we searched for missing information
independent checks of the ranking method: Semantic Scholar list of all 87 citing papers (read 2026-10-10) and full texts of PolaRiS, RoboLab, RoboWorld, PhAIL and VLA-REPLICA; web searches for critiques of RoboArena's pairwise method. Found uses of RoboArena as a reference and PhAIL's remark on confidence intervals, but no independent re-test of the ranking against another real-world measure.
live ranking model: Site code and text (calls it 'Bradley-Terry Davidson'), paper Section 3.2 and Appendix B, public repository (no ranking code). Whether the live fit uses the paper's task buckets or progress scores is not stated.
which organisations were excluded and why: Integrity notice at website commits bda3698 and a6602e6, the lead author's post, current site. The organisers name no organisations; press reports name some (secondary).
role of the 'frodobots' evaluator label: Paper, site text, transparency data, web search. No primary source describes it.
number of scenes in the live service: Paper (dozens; 32 shown), data dump card, transparency data. Not counted anywhere.
Change history
- Created at full depth from primary sources, starting from the checked basic entry. Added the June 2026 manipulation notice and rollback (site source history, lead author's post, archived official data), evaluator-concentration and self-evaluation issues from the public statistics, the leave-one-organisation reruns, the snapshot-scale anomaly, Figure 6 values, and adoption sources. The basic entry's 'official view' figures were re-read and are unchanged.
- Published as a full entry.