AutoEval

AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World

How to read this picture

AutoEval is a system that tests submitted robot policies (the models that control robots) on real robot arms without a human operator. Its public robot stations, called cells, are now offline.123

Sources
Last checked 10 Oct 2026Full entry26 of 36 facts checked at the sourceNext check 8 Apr 2027
Runs in
Real robots1
Checked against real robots
Real robots
Skill
Handling objects
Robot
One arm1
Used by
1,096 logged jobs456
63 citations
Licence
MIT7
Commercial use: allowed

What a score here does not tell you Inferred

  1. How a policy would do in other scenes or labs.The tests cover four tabletop tasks at one site.
  2. How much partial progress a policy makes, or how well it moves.A learned classifier (a model trained to judge the outcome) marks each attempt as a success or a failure.
  3. Whether a small difference in a single job is real.Public jobs run 10 episodes (attempts) by default.

Comparisons with real robots

StudyResultWhat was comparedDone by
AutoEval paper: automated compared with human-run evaluation
Mar 2025
Pearson correlation r = 0.942 and MMRV (a measure of how often two rankings disagree) = 0.015, both as means over 5 tasks1AutoEval and human operators evaluated the same 6 policies on the same real cells. The policies were OpenVLA, Octo, Open-pi0, MiniVLA, SuSIE and SuSIE's low-level policy, so there were 5 distinct models. Each policy ran 50 trials on each of 5 tasks. The study’s authors described the result as “closely match”.The benchmark’s authors

Our assessment Opinion

Automated judging matches human judges in these cells. The tasks cover a narrow range.

Reasoning

AutoEval's automation is well validated for its own cells. Automated scores closely matched human-run scores on 4 of 5 tasks. That says little about how a policy does elsewhere, because every test ran on four tabletop tasks in one lab.

Confidence: high

Single 10-episode jobs are too small to compare policies.

Reasoning

With a 10-episode job, the 95% confidence interval for a success rate near 50% spans about ±30 points (by binomial arithmetic). Use 50 episodes, as the paper did, before you compare policies.

Confidence: high

AutoEval is now mostly a method that labs can use to automate their own tests.

Reasoning

The dashboard is offline and there is no second site. AutoEval is now mainly an open method for labs that want to automate their own real-robot tests. It includes a published way to check the automation against human judges.

Confidence: medium

Known problems 4

  1. The public service is offline

    The public dashboard was offline on 2026-10-10.324

    Details

    On 2026-10-10 the dashboard address returned 'The endpoint auto-eval.ngrok.app is offline' (ERR_NGROK_3200). The site says the four public tasks are available until 2026-01-01 and that 2026 tasks are to be decided. The public logs show jobs until 2026-06-13.

  2. The public tests cover few tasks at one site

    The public service offered four tabletop tasks at one lab.128

    Details

    The public offering was 4 tabletop tasks in 2 cells at Berkeley. The authors list as limits: new scenes need hours of set-up, robustness factors such as camera angle or lighting cannot be varied in a controlled way, mobile manipulation is not covered, and success is binary. VLA-REPLICA (2026) rates AutoEval's task diversity as low (4 tasks) and its reproducibility as low.

  3. Automated judging makes errors, most often on cloth folding

    On the cloth-folding task, AutoEval counted 12 successes for one policy where humans counted 3.1

    Details

    Agreement with human judges was near perfect on the drawer and eggplant tasks but lower on cloth folding (Pearson r 0.720 by our computation). Example: Open-pi0 folded the cloth 12/50 times according to AutoEval and 3/50 according to humans. In a 50-trial failure analysis on the sink, 3 trials were wrongly counted as successes because the reset policy failed. The authors suggest humans re-check the report videos when accuracy matters.

  4. Scores drift over long runs and default jobs are small

    Scores drift after about 8 hours of continuous running. Public jobs run 10 episodes by default.19

    Details

    Scores drifted after about 8 hours of continuous running because the WidowX motors overheat; the cells pause for 20 minutes every 6 hours. The public web UI defaults to 10 episodes per job (maximum 50), while the paper used 50.

Details

About

What it is
Evaluation service Inferred12
More

Users submit a policy server; the organisers' robots run the evaluation and return a report. Classification by the Atlas.

Built by
UC Berkeley, NVIDIA1012
More

Authors: Zhiyuan Zhou, Pranav Atreya, You Liang Tan (also NVIDIA), Karl Pertsch, Sergey Levine.

Released
March 2025, at CoRL 202510114
More

arXiv v1 on 2025-03-31. Published at CoRL 2025 (PMLR volume 305, pages 1997-2017, 2025-10-07). Logged public jobs start 2025-03-12.

Code repository created 2025-03-26.

Version
0.0.1. There are no tagged releases.121310+1
More

Package auto_eval 0.0.1. No tags or releases. Paper arXiv v2; the CoRL 2025 camera-ready corrects the daily throughput in its introduction from 500 to 850 episodes.

arXiv v2's introduction says 500 episodes per 24 hours while its Section 5.3 says about 850; the camera-ready says 850 in both.

Last update
June 2026, when the last job was logged5413
More

Last logged evaluation job 2026-06-13 (dataset updated 2026-06-14). Last code commit 2026-03-26.

Status
The last code fix was in March 2026. The public service is offline. Inferred1332
More

Code fixes until March 2026; the public dashboard was offline on 2026-10-10. The site says the four public tasks run until 2026-01-01 and that 2026 tasks are to be decided.

Jobs were still logged in January (233) and March 2026 (116), then 2 in May and 1 in June 2026. See facts.access.

Setup

Runs in
Real robots1
Robot
One arm1
Robot model
WidowX 250 6-DoF arm with a Logitech C920 camera (256 x 256 top-down image)12
More

End-effector delta actions with blocking control. Policies get the image, an 8-number robot state and the instruction (site).

Setting
Tabletop1
More

Drawer, toy-sink and cloth scenes in the style of the BridgeData V2 setups.

Tasks
5 tasks (4 public) in 3 cells12
More

5 tasks in the paper (open drawer, close drawer, eggplant to basket, eggplant to sink, fold cloth) in 3 cells. 4 public tasks in 2 cells (no cloth).

Training data
No training data for the policies under test. Each cell's reset policy was trained on 50-100 teleoperated demonstrations and its success classifier on about 1,000 labelled images.1
More

Reset policies fine-tune OpenVLA with LoRA; the public cells use a scripted reset for the drawer and a fine-tuned MiniVLA for the sink (Appendix I).

Changes at test
Object positions1
More

The paper randomises the start position of the eggplant, drawer and cloth in each episode; lighting is held constant.

Scoring and access

Scored by
Success rate, Automatic judge12
More

Binary success decided by a fine-tuned PaliGemma 3B classifier, deployed only above 95% accuracy. The report gives overall_success_rate. The paper lists binary-only scoring as a limitation.

Score
Success rate, judged automatically12
More

Success rate over the job's episodes, judged automatically; results come back as a Weights & Biases report with videos, and episodes are logged to Hugging Face.

Trials
50 in the paper and 10 by default196
More

Paper · 50 rollouts per policy and task; maximum 70 steps (drawer), 100 (sink), 80 (cloth).1

Public web UI · Number of episodes defaults to 10 (allowed 1-50); steps per episode default 70. The server-side default is 50.9

World-Gymnast · 10 trials per policy and task, repeated 5 times to estimate standard error.6

Who runs it
The organisers run the tests21
More

Users submit a policy-server address; the system queues jobs and runs them on its robots.

Error bars
Sometimes reported Inferred16
More

The paper's tables give success counts out of 50 without intervals; its consistency figure shows 95% intervals. World-Gymnast reports standard errors from AutoEval runs.

Leaderboard
None. Scores are only in papers. Inferred212
More

Per-job Weights & Biases reports and a public log dataset; no ranking page.

Code licence
MIT7
More

LICENSE file, copyright 2025 zhouzypaul; README badge agrees.

Data licence
MIT for the evaluation logs on Hugging Face (dataset card)5
Access
The service is offline. The code is open. Inferred3212
More

The public submission dashboard was offline on 2026-10-10 (ngrok: endpoint offline). Code, logs and the set-up guide remain open (MIT).

The site and README still point to the dashboard. Level is our reading: the service accepts no jobs now, while anyone can build their own cell from the code.

Commercial use
Allowed Inferred75
More

Code and logs are MIT. Building a cell also involves fine-tuning PaliGemma and OpenVLA, whose own terms were not checked here. Not legal advice.

Sources 14

  1. 1AutoEval paper, full text v2 (Sections 3-5, Fig. 7, Appendices A-L, Tables 1-5)Paper · Apr 2025 · checked 10 Oct 2026
  2. 2AutoEval project site (tasks, submission instructions, FAQ)Official site · 2025 · checked 10 Oct 2026
  3. 3AutoEval dashboard URL (returned 'endpoint offline', ERR_NGROK_3200, on 2026-10-10)Official site · 10 Oct 2026 · checked 10 Oct 2026
  4. 4Hugging Face dataset viewer rows for zhouzypaul/auto_eval (job index: time, robot, location)Index · 10 Oct 2026 · checked 10 Oct 2026
  5. 5Hugging Face dataset zhouzypaul/auto_eval: card (license: mit), API record, 1,096 eval_data job foldersRepository · 14 Jun 2026 · checked 10 Oct 2026
  6. 6World-Gymnast: Training Robots with Reinforcement Learning in a World Model (Section 4, Appendix D.2)Paper · Feb 2026 · checked 10 Oct 2026
  7. 7AutoEval LICENSE (MIT)Repository · Mar 2025 · checked 10 Oct 2026
  8. 8VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of VLA Models (Table 1)Paper · May 2026 · checked 10 Oct 2026
  9. 9AutoEval web UI source (static/index.html: episodes default 10, max 50; 70 steps) and job_scheduler.pyRepository · 2025 · checked 10 Oct 2026
  10. 10AutoEval (arXiv abstract page; v1 2025-03-31, v2 2025-04-02)Paper · Mar 2025 · checked 10 Oct 2026
  11. 11AutoEval, PMLR proceedings page (CoRL 2025, PMLR 305:1997-2017)Paper · Oct 2025 · checked 10 Oct 2026
  12. 12AutoEval READMERepository · Mar 2026 · checked 10 Oct 2026
  13. 13AutoEval commit history, tags and releasesRepository · 26 Mar 2026 · checked 10 Oct 2026
  14. 14AutoEval camera-ready PDF in PMLRPaper · Oct 2025 · checked 10 Oct 2026
Where we searched for missing information

validity (independent checks of AutoEval against human-run evaluation): AutoEval arXiv v2 and camera-ready; World-Gymnast 2602.02454 (uses AutoEval, no human comparison); VLA-REPLICA 2605.20774 (characterisation only); WorldEval 2505.19017, Scalable Policy Evaluation with Video World Models 2511.11520, RoboWorld 2607.01060, RobotArena Infinity 2510.23571, PolaRiS 2512.16881 (cite only); extended web search.

access: Dashboard URL from the site and README (offline), project site, README, public logs.

industry_use: Web search, citing papers opened above, public log index.

top_score: No single headline score: per-task counts only. Skipped.

Change history

  1. Created at full depth from primary sources, starting from the basic entry and the real-eval inventory record. Added per-task agreement, camera-ready corrections, public-UI defaults, job-log counts and the World-Gymnast use.
  2. Published as a full entry.