RoboArena

RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

How to read this picture

RoboArena is a service that tests robot policies (the robots' control models) on real DROID robot arms. Volunteers run blind A/B tests, in which two unnamed policies try the same task, and their preferences are combined into a rating.123

Sources
Last checked 10 Oct 2026Full entry45 of 59 facts checked at the sourceNext check 8 Apr 2027
Runs in
Real robots2
Checked against real robots
Real robots
Skill
Handling objects
Robot
One arm2
Franka Panda (DROID setup)
Used by
Real-world reference for at least 3 evaluation papers456+2
87 citations
Licence
Commercial use: allowed

What a score here does not tell you Inferred

  1. How a policy performs on other robots.RoboArena tests only the DROID Franka setup.
  2. How often a policy succeeds on a fixed set of tasks.Scores are relative ratings from tests where evaluators choose the tasks.
  3. How good the rarely tested policies are.Of the 24 policies, 15 have fewer than 100 evaluations.

Comparisons with real robots

StudyResultWhat was comparedDone by
Oracle ranking check (paper Figure 6)
Jun 2025
Pearson r = 0.98 and MMRV = 1.8% for the task-aware model. A conventional 17-task lab protocol gave r = 0.69 and MMRV = 13%. MMRV measures how often two rankings disagree.212Rankings built from 612 A/B tests of 7 DROID policies were compared with an 'oracle' ranking. After each A/B test, the evaluator also ran the other 5 policies on the same task. The oracle ranks policies by their average progress score over 4,284 rollouts. The study’s authors described the result as “more accurately rank”.The benchmark’s authors
Simulated drift re-analysis (paper Table 2, arXiv v2)
Nov 2025
Pearson r = 0.838 and MMRV = 0.058. The conventional protocol gave r = 0.692 and MMRV = 0.141.2The same data were ranked again after simulating a shift over time from easy to hard tasks and from weaker to stronger policies. The new rankings were compared with the oracle. The study’s authors described the result as “significantly higher correlation”.The benchmark’s authors
Leave-one-organisation rerun (live audit tool)
Jun 2026
Without the largest label (57.6% of evaluations), ranks move by at most 2 places. Without 'nvidia', DreamZero falls below 100 evaluations.131415The organisers' site recomputes the official ranking with all evaluations from one evaluator organisation removed. We read the reruns for the two most influential labels on 2026-10-10.The benchmark’s authors

Our assessment Opinion

A rating shows how often evaluators preferred a policy in this pool, on this robot.

Reasoning

A RoboArena rating says how often evaluators preferred a policy over the others in the pool, on DROID Franka arms, in tasks the evaluators chose. It is a relative number that moves when the pool or the evaluators change, and it does not translate into a success rate or carry over to other robots.

Confidence: high

Check evaluation counts and evaluator shares before trusting the top places.

Reasoning

The top of the board rests on thin evidence. The official leader's place depends on evaluations carrying its developer's label, and the all-policies leader has 25 evaluations from essentially one label. Read the evaluation counts and the audit metrics next to every rating.

Confidence: high

The accuracy check uses only RoboArena's own data.

Reasoning

The paper's accuracy check compares RoboArena rankings with an oracle ranking (a reference built from the same evaluators, tasks and progress scores), over 7 related policies. It shows that the pairwise method recovers that oracle efficiently. It is not a comparison with an independent measure of real-world usefulness.

Confidence: medium

Open data made the manipulation visible. A few evaluators still supply a large share of the tests, which remains a risk.

Reasoning

Public evaluation data let the organisers and outsiders spot the spring 2026 manipulation, and the rollback can be traced in archived data. The fixes are rules and thresholds added afterwards. A large share of evaluations still comes from a few evaluators, so the board stays vulnerable to them.

Confidence: medium

Known problems 6

  1. The rollback removed 847 evaluations and changed the top of the board

    847 evaluations from April to June 2026 were removed. Spirit v1.6, which led the board before, now has 25 counted tests. Inferred161718+2

    Details

    Comparing the official data archived on 2026-06-03 with today's data: 847 of the 4,614 public evaluations then listed are gone, all dated 2026-04-02 to 2026-06-03. They came from six evaluator-organisation labels (333, 276, 142, 72, 22 and 2 evaluations); of the 954 evaluations listed for that period, 107 remain. Spirit v1.6 led with 1926 ± 37.3 from 310 evaluations; 223 of its 309 listed evaluations came from two labels that first appeared in April and May 2026 and recorded 212 wins, 1 loss and 10 ties for it, while all other evaluators recorded 33 wins, 37 losses and 16 ties. After the rollback it has 25 counted evaluations and appears only in the all-policies view. WALL-OSS (fourth on 2026-06-03, 48 evaluations) no longer appears. MolmoAct2-DROID went from 63 evaluations and 1066 to 5 evaluations and 1579. Press reports say Spirit AI had announced first place on 2026-06-03.

  2. Most evaluations come from one evaluator label

    One evaluator label supplied 57.6% of all counted evaluations and 92.9% since mid-June 2026. Inferred211813

    Details

    One evaluator-organisation label, 'frodobots', contributed 2,301 of 3,994 counted evaluations (57.6%) between 2025-08-02 and 2026-09-12. Its share was 659 of 987 in 2026 (66.8%) and 195 of 210 after 2026-06-13 (92.9%). For every official policy except DreamZero, frodobots supplied the largest share of its evaluations (53.9% to 67.5%). The official leave-one-organisation rerun without this label keeps the top two policies but moves others by up to two places.

  3. The official leader was evaluated mostly under its developer's label

    58% of the official leader's evaluations carry its developer's organisation label. Inferred182214+2

    Details

    DreamZero, an NVIDIA model, has 190 counted evaluations; 111 of them (58.4%) carry the evaluator label 'nvidia' and date from 2026-01-28 to 2026-02-26. These predate the 2026-04-02 start of the retroactive no-stake rule and remain counted. The site's own leave-one-organisation rerun without 'nvidia' leaves DreamZero with 79 evaluations, below the 100-evaluation threshold, so pi0.5-DROID would lead the official view. DreamZero has had no new evaluations since 2026-03-30 and its server was offline in the 30 days to 2026-10-11.

  4. Ratings shift with the pool, and one weekly snapshot was broken

    Scores move when other policies are added, and one weekly snapshot shifted every score by about 240 points. Inferred24

    Details

    Ratings are relative. In the published weekly snapshots, policies with no new evaluations still change score; for example Spirit v1.6 ranged from 1773 to 1792 between June and October 2026. On the 2026-08-24 snapshot every policy rose by 244 to 278 points and fell back by about 233 a week later; that snapshot gave a newly added policy with 2 evaluations a score of -3884 and a standard deviation of 39,859.4, and listed a standard deviation of about 1733 for every other policy.

  5. Evaluations were manipulated in spring 2026

    The organisers found manipulation, rolled back evaluations from April 2026 and changed the rules.232515+1

    Details

    On 2026-06-13 the RoboArena team said it had observed evidence of benchmark manipulation since April. Its notice says one sign was unusually low completion rates for requested evaluation assignments, and that organisations completing less than 20% of their requested evaluations were flagged. The team excluded evaluations from organisations with suspicious patterns and now allows only third-party evaluators with no stake in submitted policies. Both changes were applied retroactively to evaluations from 2026-04-02; the team says it found no evidence requiring exclusion before that date. It also added a 100-evaluation threshold for the official view, audit metrics and easier data downloads. The notice was removed from the site on 2026-07-17.

  6. Limitations stated by the authors

    The authors say RoboArena covers only the DROID platform and cannot easily run controlled tests that change one factor at a time. They did not study how robust it is to evaluators who act in bad faith.2

    Details

    The paper lists: evaluation on the DROID platform only; difficulty running controlled experiments that vary one factor at a time; robustness to intentionally adversarial evaluators not studied; possible over-optimisation (Goodhart's law); and possible latency from remote inference, found negligible for static tasks. It also reports that preference and progress feedback disagreed in 11% of A/B evaluations.

Details

About

What it is
Evaluation service Inferred13
More

Classified by the Atlas. Submitters host a policy server; RoboArena's evaluator network runs it on real robots. RobotArena (with the infinity sign) is a different project.

Built by
UC Berkeley, Stanford University, University of Washington, University of Montreal, NVIDIA, University of Pennsylvania, UT Austin, Yonsei University227
More

arXiv v2 lists 32 authors; the PMLR version lists 26. Eight numbered affiliations in arXiv v2.

UC Berkeley · Lead institution; two of three corresponding authors (Atreya, Pertsch).2

Stanford University · Third corresponding author (Tony Lee).2

University of Washington, University of Montreal, University of Pennsylvania, UT Austin, Yonsei University · Evaluator sites and co-authors in the paper.2

NVIDIA · Co-authors who built the simulated evaluation environments (Appendix A).2

Released
June 2025, at CoRL 202512721
More

arXiv v1 on 2025-06-22. Published at CoRL 2025 (PMLR volume 305). The first public A/B evaluation is dated 2025-04-15.

Code repository created 2025-06-14; website repository created 2025-06-17.

Version
No version numbers. The leaderboard is live, and there are 3 data dumps.282919
More

No versions, tags or releases. The leaderboard is recomputed continuously and three data dumps exist (2025-08-05, 2026-02-03, 2026-07-17).

Paper v1 (2025-06-22) and v2 (2025-11-29) · v2 adds Table 2 (simulated task and policy drift) and two authors. Figure 6 is identical in both.1212

Official and All views (since 2026-06-10) · The official view lists only policies with at least 100 A/B evaluations. The all-policies view adds entries with fewer, marked 'low sample'.315

Integrity update (2026-06-11 to 2026-06-13) · New evaluator rule, evaluations from 2026-04-02 rolled back, audit metrics added. See issues.i1.231525

Last update
The leaderboard was recomputed in October 2026. The newest evaluation is from 12 September 2026.192130+1
More

Leaderboard recomputed 2026-10-11 03:30 UTC (evening of 2026-10-10 in the US). Newest public A/B evaluation 2026-09-12. Latest data dump 2026-07-17. Code last changed 2026-04-28.

Status
Active. The site says it will run through December 2026. Inferred183
More

Evaluations every month from April to September 2026, at a lower rate than in 2025. The site says the benchmark runs live through December 2026, with possible extensions.

Counted evaluations per month in 2026 (our count from the public list): Jan 234, Feb 237, Mar 171, Apr 74, May 43, Jun 85, Jul 68, Aug 60, Sep 15 (to 2026-09-12). None between 2026-09-12 and 2026-10-10. In 2025 there were 3,008. Four of 24 policy servers were online in the 30 days to 2026-10-11.

Setup

Runs in
Real robots2
More

Closed-loop runs on physical DROID robots.

Robot
One arm2
More

One Franka arm on the DROID platform. The authors list evaluation on DROID only as a limitation.

Robot model
DROID platform: Franka Panda 7-DoF arm, Robotiq 2F-85 gripper, ZED-mini wrist camera, one or more ZED 2 external cameras, on a height-adjustable mobile table.2
Setting
Mixed Inferred2
More

Evaluators use their own labs, offices and kitchens. The paper reports dozens of scenes and shows 32 sample environments (Figure 11).

Tasks
Open-ended, with 2,797 distinct instructions Inferred218
More

No fixed task list. The paper reports hundreds of instructions; the public evaluations use 2,797 distinct instruction strings.

Distinct strings counted by us on 2026-10-10; after lower-casing and removing punctuation, 2,761. Most frequent: 'open the book' (40), 'close the book' (35).

Training data
None. Submitters use DROID data.310
More

RoboArena supplies no training data. It points submitters to the open DROID dataset and the openpi DROID training code.

Size
3,994 A/B evaluations of 24 policies21
More

Paper (2025) · 7 policies, 612 A/B comparisons, 4,284 evaluation rollouts including the oracle runs, 7 universities.2

Live service (2026-10-10) · 3,994 counted A/B evaluations (575 ties, 14.4%), 24 policies, 120 policy pairs, 60 evaluator accounts, 28 evaluator-organisation labels. First evaluation 2025-04-15, last 2026-09-12.21

Data dump 2026-07-17 · 3,883 evaluation sessions, 10,783 policy episodes, 27,148 videos, 21.7 GB (21,676,174,280 bytes).30

Removed in June 2026 · 847 evaluations dated 2026-04-02 to 2026-06-03 were taken out of the public list and the ranking (issues.i2).1718

Changes at test
Open-ended. Evaluators choose the scenes and tasks. Inferred218
More

Evaluators pick any scene, objects, camera view and instruction. Variation is wide but uncontrolled and not labelled.

Paper Figure 11: lighting, camera viewpoints, tablecloths and objects were often changed. The 3,995 public evaluations use 2,797 distinct instruction strings (our count). RoboArena does not record whether a task or object resembles the policy's training data.

Scoring and access

Scored by
Human preference, Progress score218
More

For each A/B test the evaluator gives a preference (A, B or tie in the live data), a 0 to 100 progress score per policy, and a written reason. The paper found preference and progress disagree in 11% of A/B evaluations. The live rating is described as a Bradley-Terry-Davidson fit to preferences; we did not find whether progress scores enter it.

Score
A rating from pairwise comparisons, with a standard deviation. The official view lists only policies with at least 100 tests.3192
More

Each policy gets a rating on an Elo-like scale (786 to 1788 on 2026-10-10) with a standard deviation, fitted to all counted A/B preferences, ties included. The official view lists only policies with at least 100 A/B evaluations.

The paper's preferred model adds latent task-difficulty buckets fitted by an EM algorithm (Section 3.2, Appendix B). The site calls the live method a 'Bradley-Terry Davidson ranking path' and labels its chart 'Elo score'. Ratings are relative: other policies' results move every value (issues.i5). Comparisons involving the base pi0 or pi0-FAST models are shown but not counted.

Trials
Each policy runs once in each A/B test23
More

Each A/B evaluation runs each of the two policies once, back to back, from the same start. The official view needs at least 100 A/B evaluations per policy. In the paper's oracle sessions, evaluators ran all 7 policies on each task.

Who runs it
The organisers run the tests Inferred2323
More

Step by step: (1) A submitter hosts the policy on a remote server and submits it through a form; the team checks the input and output format and runs it with a trained evaluator in a test setup before adding it to the pool. (2) An evaluator's client asks the central server for a pair and receives the two servers' IP addresses, not the policy names. (3) The evaluator arranges a scene, types an instruction and runs policy A and then policy B from the same start until success or a timeout. (4) The evaluator records a preference, two progress scores and a reason; videos and actions are uploaded. (5) The central server fits the rating. Evaluators earn evaluation credits; submitters get a weekly budget. Since June 2026 only evaluators with no stake in submitted policies may volunteer. Taxonomy fit is loose: the organisers coordinate but do not run the robots.

Error bars
Usually reported191231
More

The leaderboard shows a standard deviation for every rating, and the paper's figures carry error bars. PhAIL (2026-05) says neither RoboArena nor RoboChallenge reports confidence intervals or paired tests on the rankings.

Leaderboard
Official leaderboard193
More

Public leaderboard at robo-arena.github.io/leaderboard, with an A/B evaluation viewer and per-organisation audit metrics.

Code licence
MIT91011
More

LICENSE: 'Copyright (c) 2025 robo-arena'. It covers the public repository, which holds the policy-server interface, client example and a test script. The evaluator client, central server and ranking code are not in it (file tree checked 2026-10-10). The website repository and the linked DROID simulation repository (arhanjain/sim-evals) have no licence file.

Data licence
MIT2930
More

All three Hugging Face data dumps carry license: mit. The paper text is CC BY 4.0 on arXiv.

Asset licence
not applicable Inferred2
More

Tests use physical robots and objects; no 3D assets are distributed.

Access
Policies are submitted through a form. The data is open.32330
More

Policies are submitted through an online form and checked before they enter the pool; evaluators also sign up through a form. Results, videos and data dumps are open to all.

Commercial use
Allowed Inferred930
More

Code and data dumps are MIT-licensed. Not legal advice. The dumps contain videos recorded in evaluators' labs.

Published at
CoRL 2025 (PMLR 305:336-364)27
More

arXiv comments give no venue; the PMLR page confirms.

Sources 31

  1. 1RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv abstract page, v1 2025-06-22, v2 2025-11-29)Paper · Jun 2025 · checked 10 Oct 2026
  2. 2RoboArena paper, full text v2 (Sections 3-7, Appendices B-D)Paper · Nov 2025 · checked 10 Oct 2026
  3. 3RoboArena website and its page code (about text, leaderboard views, alias table, audit metrics)Official site · 17 Jul 2026 · checked 10 Oct 2026
  4. 4PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (v2; Figure 8, correlation with RoboArena)Paper · Dec 2025 · checked 10 Oct 2026
  5. 5RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies (v4; Fig. 10, RoboArena Elo comparison)Paper · Apr 2026 · checked 10 Oct 2026
  6. 6RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation (v4; RoboArena leaderboard as ground truth)Paper · Jul 2026 · checked 10 Oct 2026
  7. 7openpi DROID example README (RoboArena baseline checkpoints; pi0.5-DROID named strongest based on RoboArena)Repository · 2026 · checked 10 Oct 2026
  8. 8Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models (NVIDIA Technical Blog)Official blog · 15 Jun 2026 · checked 10 Oct 2026
  9. 9robo-arena/roboarena LICENSE fileRepository · Jun 2025 · checked 10 Oct 2026
  10. 10robo-arena/roboarena README and file tree (policy server interface only)Repository · 28 Apr 2026 · checked 10 Oct 2026
  11. 11GitHub API: arhanjain/sim-evals (DROID simulated evaluation linked by RoboArena; no licence)Index · 10 Oct 2026 · checked 10 Oct 2026
  12. 12RoboArena paper Figure 6: ranking accuracy versus the oracle (Pearson r and MMRV bar labels; identical in v1 and v2)Paper · Jun 2025 · checked 10 Oct 2026
  13. 13RoboArena official leave-one-organisation rerun: label 'frodobots' removedLeaderboard · Oct 2026 · checked 10 Oct 2026
  14. 14RoboArena official leave-one-organisation rerun: label 'nvidia' removedLeaderboard · Oct 2026 · checked 10 Oct 2026
  15. 15RoboArena website repository commit history (notice added 2026-06-11; 100-evaluation official view 2026-06-10; notice hidden 2026-07-17)Repository · 17 Jul 2026 · checked 10 Oct 2026
  16. 16RoboArena leaderboard data as archived on 2026-06-03 (last_updated 2026-06-03 11:49 UTC), before the rollbackLeaderboard · 3 Jun 2026 · checked 10 Oct 2026
  17. 17RoboArena public A/B evaluation list as archived on 2026-06-03 (4,614 sessions), before the rollbackLeaderboard · 3 Jun 2026 · checked 10 Oct 2026
  18. 18RoboArena public A/B evaluation list (3,995 sessions with instruction, preference, progress scores and evaluator organisation label)Leaderboard · Oct 2026 · checked 10 Oct 2026
  19. 19RoboArena leaderboard data (JSON loaded by robo-arena.github.io/leaderboard; last_updated 2026-10-11 03:30 UTC)Leaderboard · Oct 2026 · checked 10 Oct 2026
  20. 20Huxiu: report on RoboArena manipulation and Spirit v1.6 (republished from WeChat account Dianchang)Secondary · 24 Jun 2026 · checked 10 Oct 2026
  21. 21RoboArena public transparency statistics (counted A/B evaluations by policy, pair and evaluator organisation)Leaderboard · Oct 2026 · checked 10 Oct 2026
  22. 22World Action Models are Zero-shot Policies (DreamZero), arXiv 2602.15922 v1Paper · Feb 2026 · checked 10 Oct 2026
  23. 23RoboArena benchmark integrity notice, website source at commit a6602e6 ('Publish post-maintenance site')Official site · 13 Jun 2026 · checked 10 Oct 2026
  24. 24RoboArena weekly leaderboard snapshots (18 snapshots, 2026-06-10 to 2026-10-05)Leaderboard · Oct 2026 · checked 10 Oct 2026
  25. 25Post by Pranav Atreya (RoboArena lead author) on benchmark hacking and rolled-back evaluationsOfficial blog · 13 Jun 2026 · checked 10 Oct 2026
  26. 26What Do Robotics Leaderboards Tell Us About The State of Robot Learning? (It Can Think! newsletter)Secondary · 13 Jun 2026 · checked 10 Oct 2026
  27. 27RoboArena, Proceedings of the 9th Conference on Robot Learning, PMLR 305:336-364Paper · Sep 2025 · checked 10 Oct 2026
  28. 28GitHub API: robo-arena/roboarena (stars, forks, created, pushed; no releases or tags)Index · 10 Oct 2026 · checked 10 Oct 2026
  29. 29Hugging Face API: datasets by RoboArena (three data dumps, licence tags)Repository · 17 Jul 2026 · checked 10 Oct 2026
  30. 30RoboArena/DataDump_07-17-2026 dataset card and Hub record (size, downloads)Repository · 17 Jul 2026 · checked 10 Oct 2026
  31. 31PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology (Related Work)Paper · May 2026 · checked 10 Oct 2026
Where we searched for missing information

independent checks of the ranking method: Semantic Scholar list of all 87 citing papers (read 2026-10-10) and full texts of PolaRiS, RoboLab, RoboWorld, PhAIL and VLA-REPLICA; web searches for critiques of RoboArena's pairwise method. Found uses of RoboArena as a reference and PhAIL's remark on confidence intervals, but no independent re-test of the ranking against another real-world measure.

live ranking model: Site code and text (calls it 'Bradley-Terry Davidson'), paper Section 3.2 and Appendix B, public repository (no ranking code). Whether the live fit uses the paper's task buckets or progress scores is not stated.

which organisations were excluded and why: Integrity notice at website commits bda3698 and a6602e6, the lead author's post, current site. The organisers name no organisations; press reports name some (secondary).

role of the 'frodobots' evaluator label: Paper, site text, transparency data, web search. No primary source describes it.

number of scenes in the live service: Paper (dozens; 32 shown), data dump card, transparency data. Not counted anywhere.

Change history

  1. Created at full depth from primary sources, starting from the checked basic entry. Added the June 2026 manipulation notice and rollback (site source history, lead author's post, archived official data), evaluator-concentration and self-evaluation issues from the public statistics, the leave-one-organisation reruns, the snapshot-scale anomaly, Figure 6 values, and adoption sources. The basic entry's 'official view' figures were re-read and are unchanged.
  2. Published as a full entry.