RoboChallenge

RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies

How to read this picture

RoboChallenge tests robot policies (the models that control a robot) on real robots in Dexmal's lab. Users control the robots over the internet, and lab staff set up and score each test on 30 tabletop tasks.123

Sources
Last checked 10 Oct 2026Full entry47 of 61 facts checked at the sourceNext check 8 Apr 2027
Runs in
Real robots1
Checked against real robots
Real robots
Skill
Handling objects
Robot
Two arms, One arm45
ARX5, UR5, ALOHA, DOS-W1 (v2)
Used by
6 model reports678+3
34 citations
Licence
CC-BY-NC-SA-4.01213
Commercial use: not allowed

What a score here does not tell you Inferred

  1. How well a policy does in new places or with new objects.The tests use the same tasks as the training demonstrations, and each test copies the set-up of a recorded demonstration.
  2. Whether a policy would get the same score in another lab.All tests run at one site. No one has checked the results at a second site.
  3. Which model was actually tested.The model runs on the submitter's own computers, where the operator cannot check it.
  4. Whether policies got better between version 1 and version 2.Version 1 and version 2 use different tasks and rules.

Comparisons with real robots

StudyResultWhat was comparedDone by
Tester-variation test (report Figure 3)
Oct 2025
Success was 10%, 0% and 30% on pouring fries, and 50%, 70% and 80% on stacking bowls.1Two tasks were each run with one model by three kinds of tester: data collectors, first-time testers and the model's authors. The study’s authors described the result as “varies considerably”.The benchmark’s authors

Benchmarks built on RoboChallenge: CVPR 2026 RoboChallenge Track, Table 30 V2 Simulation (announced).14153

Our assessment Opinion

A score shows how well a policy repeats known tasks in one lab.

Reasoning

A Table30 score shows how well a policy, fine-tuned on RoboChallenge's own demonstrations, repeats those tasks on the same robots in the same lab. It does not show performance in new places or with new objects. The v2 zero-shot tracks, which test new objects and backgrounds, are not on the public leaderboard yet.

Confidence: high

Do not compare scores from version 1 and version 2.

Reasoning

Do not compare version 1 and version 2 numbers, and do not read the drop from 64% to 41% as a loss of skill. Tasks, robots and rules all changed between the versions.

Confidence: high

Leaderboard places are claims by the submitters, checked only by human scoring of each attempt.

Reasoning

Treat leaderboard places as claims by the submitters, checked by human scoring. The model runs on the submitter's side, entry names can be anonymous, and the newest run counts. The operator's own model family is near the top of both versions.

Confidence: medium

Do not read much into gaps of a few points. Check the date and track behind any first-place claim.

Reasoning

With 10 attempts per task and no error bars, gaps of a few points between entries are within the normal variation between runs. Many first-place claims refer to different dates or tracks.

Confidence: medium

Known problems 7

  1. The same entries have different numbers in different sources

    The RoboChallenge report, the leaderboard and the model makers' own reports give different numbers for the same entries.11617+6

    Details

    The report's pi0.5 baseline is 43.7% / 62.2; the leaderboard shows 42.67% / 61.84, with two tasks different (arrange fruits 80% to 40%, sort electronic products 40% to 50%). pi0's progress score is 47.6 in the report and 46.41 on the board. Dexmal's OpenDM README gives VLA-DM0.5 43.0% success; the board shows 40.67% with the same 54.42 score. Qwen's report cites DM0_generalist at 48.43; the board shows 49.08. The V2 paper and the site differ on tasks per robot. In December 2025 the site warned that some displayed results were temporary, partial or for debugging.

  2. The operator's own models compete on its leaderboard

    Dexmal, which runs and scores the tests, also has models near the top of the leaderboard. Inferred16186+3

    Details

    Dexmal staff run the robots, set-ups and scoring, and Dexmal's DM0 ranks second on the final v1 board (62.00%); DM0_generalist is eighth. The current v2 leader, VLA-DM0.5, is a fine-tune of Dexmal's DM0.5, whose Table30 v2 checkpoints Dexmal released in August 2026. Several committee partners also have entries.

  3. Version 2 ranking rules reward re-runs and rank incomplete entries

    In version 2, only the newest run of each task counts, so a team can re-run a task until it is satisfied. Entries with untested tasks are still ranked. Inferred182122+2

    Details

    In v2, the latest ranked run of each task replaces earlier ones (all 46 tasks with several ranked runs match the latest run, only 21 match the best), so a team can re-run a task and stop when satisfied. Untested tasks count as 0 and partial entries are ranked (second place has 20 of 30 tasks). v1's overall board required all 30 tasks.

  4. Scores from 10 attempts per task are shown without error bars

    Each score is a single number from 10 attempts per task. No measure of uncertainty is shown.241821

    Details

    Scores rest on 10 rollouts per task and are published without intervals. PhAIL notes that neither RoboArena nor RoboChallenge reports confidence intervals or paired tests. Repeated ranked v2 runs of the same entry and task differ by a median of 10 points (our count).

  5. The operator cannot check which model ran

    Models run on the submitters' own computers. The organisers cannot verify which model was actually used.1202+1

    Details

    Inference runs on the submitter's computers. The report says the platform has no means to check that the model actually run matches the claimed one; a user could use per-task models when a generalist is expected, or even intervene by hand. The annual report repeats this. For v2 the operator added a timing check (the model must declare it is ready before it learns the task) and calls it necessary but not sufficient; enforcement rests on users accepting the rules. Entries can use anonymous names: Qwen submitted as 'Lira_generalist'.

  6. The person who sets up the test changes the score

    In the RoboChallenge report's study, changing the person who set up the test changed one model's success rate by up to 30 points.1202

    Details

    In the report's tester study, the same model scored 10%, 0% and 30% on 'pour french fries' and 50%, 70% and 80% on 'stack bowls' with experienced, first-time and adaptive (model-author) testers. Adaptive testers placed objects in 'sweet spots'. The annual report says lighting and the tester's clothing changed a model's score between two submissions. The V2 paper says camera and arm parameters differ between robot units.

  7. Licences do not match the stated policy

    The committee's policy says data and code are published under open licences. The datasets have no licence, and the official client software is licensed for non-commercial use only. Inferred252612

    Details

    The committee's governance model says results, datasets and code are to be published under open licences. The datasets carry no licence and the official client is non-commercial (CC BY-NC-SA 4.0).

Details

About

What it is
Evaluation service Inferred123
More

Classified by the Atlas. RoboChallenge runs the robots and scores the runs; the submitter runs the model on its own computers over the internet.

Built by
Dexmal, Hugging Face12027
More

Report authors are listed in alphabetical order.

Dexmal (原力灵机) · 35 of 37 report authors and all 7 V2-paper authors. Supplies the robot cluster and staff. Legal name 北京原力灵机智能科技有限公司 (Beijing).1227+1

Hugging Face · 2 of 37 report authors; named co-initiator.120

RoboChallenge Committee (from 2025-11-20) · Partners named in the annual report: BAAI, AgiBot, Qwen, Galaxea, X Square Robot, Tsinghua University, Xi'an Jiaotong University, GOSIM.202825+1

Released
October 202532030
More

Site online 2025-10-17 (news item); the annual report gives 2025-10-15 as launch; arXiv report 2025-10-20.

Version
Table30 v2 (2026). Version 1 was retired in May 2026. Scores from the two versions are not comparable.543+1
More

Two benchmark versions with separate leaderboards. Table30 (v1.0) ran from October 2025 and was retired on 2026-05-27. Table30 v2 (v2.0) opened as a preview for the CVPR 2026 competition in April 2026; its task data is dated 2026-05-28.

Table30 v1.0 · 30 tasks on ALOHA (11), ARX5 (11), UR5 (6) and Franka (2); 19 single-arm and 11 two-arm. Task specs dated 2025-09-24. Retired 2026-05-27 after 'over 80,000 real-world robot task executions' and '20+ models'.53

Table30 v2.0 · 30 tasks on DOS-W1 (10), ALOHA (10), ARX5 (7) and UR5 (3); 20 dual-arm. 18 new tasks plus 12 carried from v1 in revised form; Franka removed; multi-task models only; set-up alignment relaxed.4231

v1 and v2 scores are not comparable · No task is identical across versions: 3 task names recur with changed props, steps or robot (for example three flowers in v1, four in v2). Robots, ranking rules and set-up rules also changed. We found no conversion between the two scales.542+1

Table 30 V2 Simulation (announced) · 2026-08-19: a simulation copy of Table30 v2 in NVIDIA Isaac Lab, with Lightwheel and NVIDIA, aligned in tasks and scoring with the real benchmark. Not released as of 2026-10-10.315

Last update
October 202621183
More

Newest Table30 v2 run ended 2026-10-10; leaderboard files regenerated 2026-10-10 07:00 GMT. Last news item 2026-08-19.

Status
Active. Table30 v2 has had public runs every month since April 2026. Inferred21
More

Table30 v2 runs every month since April 2026 (141, 298, 41, 162, 71, 289 and 17 public runs from April to October).

Setup

Runs in
Real robots1
Robot
Two arms, One arm45
More

v2: 20 dual-arm and 10 single-arm tasks. v1: 11 two-arm and 19 single-arm.

Robot model
v2: ARX5, UR5, Cobot Magic ALOHA and DOS-W1 (a mobile two-arm robot with Airbot Play arms). v1: UR5 with Robotiq gripper, Franka (Robotiq gripper), Cobot Magic ALOHA and ARX-5. Intel RealSense cameras.1232
More

10 machines at launch (report); 20 by January 2026 (annual report).

Setting
Tabletop1
More

'All tasks are executed on the table, or around a table.' One testing site.

Tasks
30 tasks in each version54
More

30 tasks in each version; 3 task names recur, with changes.

Training data
32,980 demonstrations in v24533+1
More

v2: 32,980 demonstration episodes (1,001 to 1,675 per task), about 1.55 TB. v1: up to 1,000 per task; the site lists 25,627 episodes, about 1.57 TB.

The v2 total equals the sum of per-task counts. Reading v1's episode_number field as demonstrations is ours.

Size
About 50,000 public test attempts across both versions Inferred3421
More

Our counts. The public run lists hold fewer rollouts than the platform totals, so they appear to be a subset.

v1 platform totals · 41,969 rollouts, 209 test tokens issued and 82 developers submitting by January 2026 (annual report); 'over 80,000' robot task executions by 2026-05-27 (news).203

v1 public run list · 3,985 runs, 39,852 rollouts, 125 user names, 2025-09-20 to 2026-05-25.34

v2 public run list · 1,019 ranked runs, 10,189 rollouts, 51 user names, 2026-04-13 to 2026-10-10.21

Leaderboards · v1 final: 22 entries, all with 30 tasks. v2: 53 entries, 2 with all 30 tasks.1618

Changes at test
Only the start positions change. Each test set-up copies a held-out demonstration. Inferred1218
More

Test set-ups copy held-out demonstration episodes. v1 used a pixel-level image overlay; v2 relaxes this to rough alignment and assigns a random robot unit of the right type. Zero-shot object and background tracks are described for v2, but no such scores appear on the public board.

The V2 paper says its competition preview does not score zero-shot or out-of-domain settings and that the out-of-domain mode will not be ranked.

Scoring and access

Scored by
Success rate, Progress score12022
More

Ranked by success rate, then progress score.

Score
Average success rate over 30 tasks. A progress score breaks ties.12022
More

Per task: success rate over 10 rollouts and a progress score (10 stage points per rollout, minus 0.5 per retry, so 0 to 100). Overall: plain averages over the 30 tasks; ranking by success rate, then progress score.

v1's overall board lists only entries that completed all 30 tasks. v2 averages over all 30 tasks with untested tasks counted as 0, ranks partial entries, and uses the latest ranked run per task (issues.i5). The V2 paper announces rollout timeouts and a 'time to complete' score; no time field appears in the v2 leaderboard data on 2026-10-10.

Trials
10 attempts per task12021
More

10 rollouts per task per run; a full entry is 300 rollouts.

Who runs it
The organisers run the tests Inferred1202+2
More

Step by step: (1) Apply for a test token; download the task data from Hugging Face and fine-tune. (2) Submit a request naming the model and tasks; in v2 a ranked run must cover a whole robot group, and single-task 'test' runs are not ranked. (3) Staff queue the job (hours to days) and tell the user when to have the model running; in v2 the model must declare it is ready before it learns which task comes next. (4) For each rollout a tester places the props to match a reference image from a held-out demonstration and watches the run. (5) The user's program pulls camera images and robot state through RoboChallenge's API and pushes actions into a queue on the robot. (6) Testers score each stage by hand; scores get a second review and can be appealed. (7) Videos and logs are published.

Error bars
Not reported Inferred181624
More

Leaderboard data give success rate and score only. PhAIL (2026-05) notes RoboChallenge reports no confidence intervals or paired tests.

Leaderboard
Official leaderboard1816
Code licence
CC-BY-NC-SA-4.01213
More

RoboChallengeInference (client and mock server) LICENSE text is CC BY-NC-SA 4.0; GitHub's API shows NOASSERTION. The organisation's openpi fork is Apache-2.0; the Community repo has no licence. Dexmal's separate OpenDM client is Apache-2.0.

Data licence
Unknown333226+1
More

No licence field or text in the Table30 and Table30v2 dataset cards; none of the 63 RoboChallenge datasets on Hugging Face has a licence tag (checked 2026-10-10). The governance document says datasets are to be published under open licences.

Asset licence
not applicable Inferred1
More

Physical robots and props; no 3D assets distributed.

Access
Testing needs an approved application. The data is open to download.202333
More

Evaluation needs an approved account and test token (209 issued by January 2026). Data downloads are open, without registration.

Commercial use
Not allowed Inferred1226
More

The official client is CC BY-NC-SA 4.0 and the data has no stated licence. Users can write their own client from the documented API. Not legal advice.

Published at
Technical report arXiv 2510.17950 (single version). Table30 V2 paper: CVPR 2026 Workshops (GigaBrain Challenge), pp. 4461-4467.302

Sources 34

  1. 1RoboChallenge technical report, full text v1 (Sections 2-4, Appendix A; Figure 3 tester study)Paper · Oct 2025 · checked 10 Oct 2026
  2. 2Table30 V2: Evaluating Generalized Models by Real Robots at Scale (CVPR 2026 Workshops, pp. 4461-4467)Paper · Jun 2026 · checked 10 Oct 2026
  3. 3RoboChallenge news page (dated items 2025-10-17 to 2026-08-19; read from the site's page code)Official site · 19 Aug 2026 · checked 10 Oct 2026
  4. 4Table 30 V2 (v2.0) benchmark metadata: tasks, robots, prompts, scoring stages, episode countsLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
  5. 5Table 30 (v1.0) benchmark metadata: tasks, robots, scoring stagesLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
  6. 6DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI (Section 4.2, RoboChallenge results)Paper · Feb 2026 · checked 10 Oct 2026
  7. 7Qwen-RobotManip Technical Report (v2; Section 6.3.2, Table30-v1 generalist track)Paper · Jun 2026 · checked 10 Oct 2026
  8. 8StarVLA-alpha: Reducing Complexity in Vision-Language-Action Systems (v2; Section 5, Appendix E)Paper · Apr 2026 · checked 10 Oct 2026
  9. 9Spirit-v1.5 repository README ('ranks #1 on RoboChallenge Table30' as of 2026-01-11)Repository · Jan 2026 · checked 10 Oct 2026
  10. 10GigaBrain-0 repository README (news: GigaBrain-0.1 first place on RoboChallenge, 2026-02-09)Repository · Aug 2026 · checked 10 Oct 2026
  11. 11Dexmal OpenDM README (news 2026-10-05: KDDI Research's VLA-DM0.5 tops Table30 V2; results table)Repository · 5 Oct 2026 · checked 10 Oct 2026
  12. 12RoboChallengeInference LICENSE (CC BY-NC-SA 4.0) and READMERepository · Oct 2025 · checked 10 Oct 2026
  13. 13GitHub API: RoboChallenge organisation repositories (stars, licences)Index · 10 Oct 2026 · checked 10 Oct 2026
  14. 14RoboChallenge CVPR 2026 competition page and task list (same 30 task IDs as v2)Official site · Apr 2026 · checked 10 Oct 2026
  15. 15RoboChallenge Table 30 V2 simulation page (Lightwheel, NVIDIA Isaac Lab)Official site · 19 Aug 2026 · checked 10 Oct 2026
  16. 16Table 30 (v1) final leaderboard data (22 entries)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
  17. 17Table 30 (v1) per-task leaderboard dataLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
  18. 18Table 30 V2 leaderboard data (53 entries, per-task results)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
  19. 19Dexmal/DM05-Table30v2-W1 model card (DM0.5 checkpoints for Table30 v2)Repository · 6 Aug 2026 · checked 10 Oct 2026
  20. 202025 RoboChallenge 年度报告 (Annual Report 2025 Q4 - 2026 Q1, Chinese edition)Report · 30 Jan 2026 · checked 10 Oct 2026
  21. 21Table 30 V2 public run list (1,019 ranked runs with rollout scores)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
  22. 22RoboChallenge leaderboard page code (ranking order; v2 entries treated as multi-task)Official site · 10 Oct 2026 · checked 10 Oct 2026
  23. 23RoboChallenge v2 submission form code (ranked versus test runs; whole robot group required)Official site · 10 Oct 2026 · checked 10 Oct 2026
  24. 24PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology (Related Work)Paper · May 2026 · checked 10 Oct 2026
  25. 25RoboChallenge Committee Governance ModelOfficial site · Nov 2025 · checked 10 Oct 2026
  26. 26Hugging Face API: all 63 RoboChallenge datasets (no licence tags)Repository · Apr 2026 · checked 10 Oct 2026
  27. 27Dexmal official website (北京原力灵机智能科技有限公司; robots and models)Official site · Oct 2026 · checked 10 Oct 2026
  28. 282025 RoboChallenge Annual Report (English edition)Report · 30 Jan 2026 · checked 10 Oct 2026
  29. 29RoboChallenge Community README (Chinese edition)Repository · 30 Jan 2026 · checked 10 Oct 2026
  30. 30RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies (arXiv abstract page, single version v1)Paper · Oct 2025 · checked 10 Oct 2026
  31. 31QbitAI: authorised reprint of RoboChallenge's Table30 V2 launch announcementSecondary · 24 Mar 2026 · checked 10 Oct 2026
  32. 32RoboChallenge/Table30v2 dataset card and Hub recordRepository · 5 Jun 2026 · checked 10 Oct 2026
  33. 33RoboChallenge/Table30 dataset card and Hub recordRepository · 13 Jan 2026 · checked 10 Oct 2026
  34. 34Table 30 (v1) public run list (3,985 runs with rollout scores)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information

license_data: Table30 and Table30v2 dataset cards, Hugging Face API for all 63 RoboChallenge datasets, governance PDF, annual reports, site code.

cross-site or independent checks: Report, V2 paper, annual reports (Chinese and English), Semantic Scholar list of 34 citing papers (full texts of 10 scanned), web search for critiques in English and Chinese. Found only PhAIL's remark on uncertainty.

zero-shot and time-to-complete scores: v2 leaderboard and run data, benchmark metadata, site code. Not present on 2026-10-10.

official Chinese announcement of v2: Site news, Community README, web search for the WeChat original; only the QbitAI reprint (secondary) was found. The CVPR 2026 Workshops paper is used instead.

statement on v1/v2 comparability: V2 paper, site news, benchmark metadata. No explicit statement; our conclusion is inferred.

Change history

  1. Created at full depth from primary sources, starting from the checked basic entry. Added the Table30 V2 CVPRW paper, both annual reports, governance model, v1/v2 comparison, leaderboard rules from site code and data, run counts, model reports and tester study. Resolved the basic entry's issue on Qwen-RobotManip: its 'ranks 1st' claim refers to the v1 generalist track, where it is listed as Lira_generalist.
  2. Published as a full entry.