AutoEval
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
AutoEval is a system that tests submitted robot policies (the models that control robots) on real robot arms without a human operator. Its public robot stations, called cells, are now offline.123
What a score here does not tell you Inferred
- How a policy would do in other scenes or labs.The tests cover four tabletop tasks at one site.
- How much partial progress a policy makes, or how well it moves.A learned classifier (a model trained to judge the outcome) marks each attempt as a success or a failure.
- Whether a small difference in a single job is real.Public jobs run 10 episodes (attempts) by default.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| AutoEval paper: automated compared with human-run evaluation Mar 2025 | Pearson correlation r = 0.942 and MMRV (a measure of how often two rankings disagree) = 0.015, both as means over 5 tasks1 | AutoEval and human operators evaluated the same 6 policies on the same real cells. The policies were OpenVLA, Octo, Open-pi0, MiniVLA, SuSIE and SuSIE's low-level policy, so there were 5 distinct models. Each policy ran 50 trials on each of 5 tasks. The study’s authors described the result as “closely match”. | The benchmark’s authors |
Our assessment Opinion
Automated judging matches human judges in these cells. The tasks cover a narrow range.
Reasoning
AutoEval's automation is well validated for its own cells. Automated scores closely matched human-run scores on 4 of 5 tasks. That says little about how a policy does elsewhere, because every test ran on four tabletop tasks in one lab.
Confidence: high
Single 10-episode jobs are too small to compare policies.
Reasoning
With a 10-episode job, the 95% confidence interval for a success rate near 50% spans about ±30 points (by binomial arithmetic). Use 50 episodes, as the paper did, before you compare policies.
Confidence: high
AutoEval is now mostly a method that labs can use to automate their own tests.
Reasoning
The dashboard is offline and there is no second site. AutoEval is now mainly an open method for labs that want to automate their own real-robot tests. It includes a published way to check the automation against human judges.
Confidence: medium
Known problems 4
The public service is offline
The public dashboard was offline on 2026-10-10.324
Details
On 2026-10-10 the dashboard address returned 'The endpoint auto-eval.ngrok.app is offline' (ERR_NGROK_3200). The site says the four public tasks are available until 2026-01-01 and that 2026 tasks are to be decided. The public logs show jobs until 2026-06-13.
The public tests cover few tasks at one site
The public service offered four tabletop tasks at one lab.128
Details
The public offering was 4 tabletop tasks in 2 cells at Berkeley. The authors list as limits: new scenes need hours of set-up, robustness factors such as camera angle or lighting cannot be varied in a controlled way, mobile manipulation is not covered, and success is binary. VLA-REPLICA (2026) rates AutoEval's task diversity as low (4 tasks) and its reproducibility as low.
Automated judging makes errors, most often on cloth folding
On the cloth-folding task, AutoEval counted 12 successes for one policy where humans counted 3.1
Details
Agreement with human judges was near perfect on the drawer and eggplant tasks but lower on cloth folding (Pearson r 0.720 by our computation). Example: Open-pi0 folded the cloth 12/50 times according to AutoEval and 3/50 according to humans. In a 50-trial failure analysis on the sink, 3 trials were wrongly counted as successes because the reset policy failed. The authors suggest humans re-check the report videos when accuracy matters.
Scores drift over long runs and default jobs are small
Scores drift after about 8 hours of continuous running. Public jobs run 10 episodes by default.19
Details
Scores drifted after about 8 hours of continuous running because the WidowX motors overheat; the cells pause for 20 minutes every 6 hours. The public web UI defaults to 10 episodes per job (maximum 50), while the paper used 50.
Details
About
- What it is
- Evaluation service Inferred12
More
Users submit a policy server; the organisers' robots run the evaluation and return a report. Classification by the Atlas.
- Built by
- UC Berkeley, NVIDIA1012
More
Authors: Zhiyuan Zhou, Pranav Atreya, You Liang Tan (also NVIDIA), Karl Pertsch, Sergey Levine.
- Released
- March 2025, at CoRL 202510114
More
arXiv v1 on 2025-03-31. Published at CoRL 2025 (PMLR volume 305, pages 1997-2017, 2025-10-07). Logged public jobs start 2025-03-12.
Code repository created 2025-03-26.
- Version
- 0.0.1. There are no tagged releases.121310+1
More
Package auto_eval 0.0.1. No tags or releases. Paper arXiv v2; the CoRL 2025 camera-ready corrects the daily throughput in its introduction from 500 to 850 episodes.
arXiv v2's introduction says 500 episodes per 24 hours while its Section 5.3 says about 850; the camera-ready says 850 in both.
- Last update
- June 2026, when the last job was logged5413
More
Last logged evaluation job 2026-06-13 (dataset updated 2026-06-14). Last code commit 2026-03-26.
- Status
- The last code fix was in March 2026. The public service is offline. Inferred1332
More
Code fixes until March 2026; the public dashboard was offline on 2026-10-10. The site says the four public tasks run until 2026-01-01 and that 2026 tasks are to be decided.
Jobs were still logged in January (233) and March 2026 (116), then 2 in May and 1 in June 2026. See facts.access.
Setup
- Runs in
- Real robots1
- Robot
- One arm1
- Robot model
- WidowX 250 6-DoF arm with a Logitech C920 camera (256 x 256 top-down image)12
More
End-effector delta actions with blocking control. Policies get the image, an 8-number robot state and the instruction (site).
- Setting
- Tabletop1
More
Drawer, toy-sink and cloth scenes in the style of the BridgeData V2 setups.
- Tasks
- 5 tasks (4 public) in 3 cells12
More
5 tasks in the paper (open drawer, close drawer, eggplant to basket, eggplant to sink, fold cloth) in 3 cells. 4 public tasks in 2 cells (no cloth).
- Training data
- No training data for the policies under test. Each cell's reset policy was trained on 50-100 teleoperated demonstrations and its success classifier on about 1,000 labelled images.1
More
Reset policies fine-tune OpenVLA with LoRA; the public cells use a scripted reset for the drawer and a fine-tuned MiniVLA for the sink (Appendix I).
- Changes at test
- Object positions1
More
The paper randomises the start position of the eggplant, drawer and cloth in each episode; lighting is held constant.
Scoring and access
- Scored by
- Success rate, Automatic judge12
More
Binary success decided by a fine-tuned PaliGemma 3B classifier, deployed only above 95% accuracy. The report gives overall_success_rate. The paper lists binary-only scoring as a limitation.
- Score
- Success rate, judged automatically12
More
Success rate over the job's episodes, judged automatically; results come back as a Weights & Biases report with videos, and episodes are logged to Hugging Face.
- Trials
- 50 in the paper and 10 by default196
More
Paper · 50 rollouts per policy and task; maximum 70 steps (drawer), 100 (sink), 80 (cloth).1
Public web UI · Number of episodes defaults to 10 (allowed 1-50); steps per episode default 70. The server-side default is 50.9
World-Gymnast · 10 trials per policy and task, repeated 5 times to estimate standard error.6
- Who runs it
- The organisers run the tests21
More
Users submit a policy-server address; the system queues jobs and runs them on its robots.
- Error bars
- Sometimes reported Inferred16
More
The paper's tables give success counts out of 50 without intervals; its consistency figure shows 95% intervals. World-Gymnast reports standard errors from AutoEval runs.
- Leaderboard
- None. Scores are only in papers. Inferred212
More
Per-job Weights & Biases reports and a public log dataset; no ranking page.
- Code licence
- MIT7
More
LICENSE file, copyright 2025 zhouzypaul; README badge agrees.
- Data licence
- MIT for the evaluation logs on Hugging Face (dataset card)5
- Access
- The service is offline. The code is open. Inferred3212
More
The public submission dashboard was offline on 2026-10-10 (ngrok: endpoint offline). Code, logs and the set-up guide remain open (MIT).
The site and README still point to the dashboard. Level is our reading: the service accepts no jobs now, while anyone can build their own cell from the code.
Sources 14
- 1AutoEval paper, full text v2 (Sections 3-5, Fig. 7, Appendices A-L, Tables 1-5)Paper · Apr 2025 · checked 10 Oct 2026
- 2AutoEval project site (tasks, submission instructions, FAQ)Official site · 2025 · checked 10 Oct 2026
- 3AutoEval dashboard URL (returned 'endpoint offline', ERR_NGROK_3200, on 2026-10-10)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 4Hugging Face dataset viewer rows for zhouzypaul/auto_eval (job index: time, robot, location)Index · 10 Oct 2026 · checked 10 Oct 2026
- 5Hugging Face dataset zhouzypaul/auto_eval: card (license: mit), API record, 1,096 eval_data job foldersRepository · 14 Jun 2026 · checked 10 Oct 2026
- 6World-Gymnast: Training Robots with Reinforcement Learning in a World Model (Section 4, Appendix D.2)Paper · Feb 2026 · checked 10 Oct 2026
- 7AutoEval LICENSE (MIT)Repository · Mar 2025 · checked 10 Oct 2026
- 8VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of VLA Models (Table 1)Paper · May 2026 · checked 10 Oct 2026
- 9AutoEval web UI source (static/index.html: episodes default 10, max 50; 70 steps) and job_scheduler.pyRepository · 2025 · checked 10 Oct 2026
- 10AutoEval (arXiv abstract page; v1 2025-03-31, v2 2025-04-02)Paper · Mar 2025 · checked 10 Oct 2026
- 11AutoEval, PMLR proceedings page (CoRL 2025, PMLR 305:1997-2017)Paper · Oct 2025 · checked 10 Oct 2026
- 12AutoEval READMERepository · Mar 2026 · checked 10 Oct 2026
- 13AutoEval commit history, tags and releasesRepository · 26 Mar 2026 · checked 10 Oct 2026
- 14AutoEval camera-ready PDF in PMLRPaper · Oct 2025 · checked 10 Oct 2026
Where we searched for missing information
validity (independent checks of AutoEval against human-run evaluation): AutoEval arXiv v2 and camera-ready; World-Gymnast 2602.02454 (uses AutoEval, no human comparison); VLA-REPLICA 2605.20774 (characterisation only); WorldEval 2505.19017, Scalable Policy Evaluation with Video World Models 2511.11520, RoboWorld 2607.01060, RobotArena Infinity 2510.23571, PolaRiS 2512.16881 (cite only); extended web search.
access: Dashboard URL from the site and README (offline), project site, README, public logs.
industry_use: Web search, citing papers opened above, public log index.
top_score: No single headline score: per-task counts only. Skipped.
Change history
- Created at full depth from primary sources, starting from the basic entry and the real-eval inventory record. Added per-task agreement, camera-ready corrections, public-UI defaults, job-log counts and the World-Gymnast use.
- Published as a full entry.