RoboChallenge
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
RoboChallenge tests robot policies (the models that control a robot) on real robots in Dexmal's lab. Users control the robots over the internet, and lab staff set up and score each test on 30 tabletop tasks.123
What a score here does not tell you Inferred
- How well a policy does in new places or with new objects.The tests use the same tasks as the training demonstrations, and each test copies the set-up of a recorded demonstration.
- Whether a policy would get the same score in another lab.All tests run at one site. No one has checked the results at a second site.
- Which model was actually tested.The model runs on the submitter's own computers, where the operator cannot check it.
- Whether policies got better between version 1 and version 2.Version 1 and version 2 use different tasks and rules.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| Tester-variation test (report Figure 3) Oct 2025 | Success was 10%, 0% and 30% on pouring fries, and 50%, 70% and 80% on stacking bowls.1 | Two tasks were each run with one model by three kinds of tester: data collectors, first-time testers and the model's authors. The study’s authors described the result as “varies considerably”. | The benchmark’s authors |
Benchmarks built on RoboChallenge: CVPR 2026 RoboChallenge Track, Table 30 V2 Simulation (announced).14153
Our assessment Opinion
A score shows how well a policy repeats known tasks in one lab.
Reasoning
A Table30 score shows how well a policy, fine-tuned on RoboChallenge's own demonstrations, repeats those tasks on the same robots in the same lab. It does not show performance in new places or with new objects. The v2 zero-shot tracks, which test new objects and backgrounds, are not on the public leaderboard yet.
Confidence: high
Do not compare scores from version 1 and version 2.
Reasoning
Do not compare version 1 and version 2 numbers, and do not read the drop from 64% to 41% as a loss of skill. Tasks, robots and rules all changed between the versions.
Confidence: high
Leaderboard places are claims by the submitters, checked only by human scoring of each attempt.
Reasoning
Treat leaderboard places as claims by the submitters, checked by human scoring. The model runs on the submitter's side, entry names can be anonymous, and the newest run counts. The operator's own model family is near the top of both versions.
Confidence: medium
Do not read much into gaps of a few points. Check the date and track behind any first-place claim.
Reasoning
With 10 attempts per task and no error bars, gaps of a few points between entries are within the normal variation between runs. Many first-place claims refer to different dates or tracks.
Confidence: medium
Known problems 7
The same entries have different numbers in different sources
The RoboChallenge report, the leaderboard and the model makers' own reports give different numbers for the same entries.11617+6
Details
The report's pi0.5 baseline is 43.7% / 62.2; the leaderboard shows 42.67% / 61.84, with two tasks different (arrange fruits 80% to 40%, sort electronic products 40% to 50%). pi0's progress score is 47.6 in the report and 46.41 on the board. Dexmal's OpenDM README gives VLA-DM0.5 43.0% success; the board shows 40.67% with the same 54.42 score. Qwen's report cites DM0_generalist at 48.43; the board shows 49.08. The V2 paper and the site differ on tasks per robot. In December 2025 the site warned that some displayed results were temporary, partial or for debugging.
The operator's own models compete on its leaderboard
Dexmal, which runs and scores the tests, also has models near the top of the leaderboard. Inferred16186+3
Details
Dexmal staff run the robots, set-ups and scoring, and Dexmal's DM0 ranks second on the final v1 board (62.00%); DM0_generalist is eighth. The current v2 leader, VLA-DM0.5, is a fine-tune of Dexmal's DM0.5, whose Table30 v2 checkpoints Dexmal released in August 2026. Several committee partners also have entries.
Version 2 ranking rules reward re-runs and rank incomplete entries
In version 2, only the newest run of each task counts, so a team can re-run a task until it is satisfied. Entries with untested tasks are still ranked. Inferred182122+2
Details
In v2, the latest ranked run of each task replaces earlier ones (all 46 tasks with several ranked runs match the latest run, only 21 match the best), so a team can re-run a task and stop when satisfied. Untested tasks count as 0 and partial entries are ranked (second place has 20 of 30 tasks). v1's overall board required all 30 tasks.
Scores from 10 attempts per task are shown without error bars
Each score is a single number from 10 attempts per task. No measure of uncertainty is shown.241821
Details
Scores rest on 10 rollouts per task and are published without intervals. PhAIL notes that neither RoboArena nor RoboChallenge reports confidence intervals or paired tests. Repeated ranked v2 runs of the same entry and task differ by a median of 10 points (our count).
The operator cannot check which model ran
Models run on the submitters' own computers. The organisers cannot verify which model was actually used.1202+1
Details
Inference runs on the submitter's computers. The report says the platform has no means to check that the model actually run matches the claimed one; a user could use per-task models when a generalist is expected, or even intervene by hand. The annual report repeats this. For v2 the operator added a timing check (the model must declare it is ready before it learns the task) and calls it necessary but not sufficient; enforcement rests on users accepting the rules. Entries can use anonymous names: Qwen submitted as 'Lira_generalist'.
The person who sets up the test changes the score
In the RoboChallenge report's study, changing the person who set up the test changed one model's success rate by up to 30 points.1202
Details
In the report's tester study, the same model scored 10%, 0% and 30% on 'pour french fries' and 50%, 70% and 80% on 'stack bowls' with experienced, first-time and adaptive (model-author) testers. Adaptive testers placed objects in 'sweet spots'. The annual report says lighting and the tester's clothing changed a model's score between two submissions. The V2 paper says camera and arm parameters differ between robot units.
Licences do not match the stated policy
The committee's policy says data and code are published under open licences. The datasets have no licence, and the official client software is licensed for non-commercial use only. Inferred252612
Details
The committee's governance model says results, datasets and code are to be published under open licences. The datasets carry no licence and the official client is non-commercial (CC BY-NC-SA 4.0).
Details
About
- What it is
- Evaluation service Inferred123
More
Classified by the Atlas. RoboChallenge runs the robots and scores the runs; the submitter runs the model on its own computers over the internet.
- Built by
- Dexmal, Hugging Face12027
More
Report authors are listed in alphabetical order.
Dexmal (原力灵机) · 35 of 37 report authors and all 7 V2-paper authors. Supplies the robot cluster and staff. Legal name 北京原力灵机智能科技有限公司 (Beijing).1227+1
Hugging Face · 2 of 37 report authors; named co-initiator.120
RoboChallenge Committee (from 2025-11-20) · Partners named in the annual report: BAAI, AgiBot, Qwen, Galaxea, X Square Robot, Tsinghua University, Xi'an Jiaotong University, GOSIM.202825+1
- Released
- October 202532030
More
Site online 2025-10-17 (news item); the annual report gives 2025-10-15 as launch; arXiv report 2025-10-20.
- Version
- Table30 v2 (2026). Version 1 was retired in May 2026. Scores from the two versions are not comparable.543+1
More
Two benchmark versions with separate leaderboards. Table30 (v1.0) ran from October 2025 and was retired on 2026-05-27. Table30 v2 (v2.0) opened as a preview for the CVPR 2026 competition in April 2026; its task data is dated 2026-05-28.
Table30 v1.0 · 30 tasks on ALOHA (11), ARX5 (11), UR5 (6) and Franka (2); 19 single-arm and 11 two-arm. Task specs dated 2025-09-24. Retired 2026-05-27 after 'over 80,000 real-world robot task executions' and '20+ models'.53
Table30 v2.0 · 30 tasks on DOS-W1 (10), ALOHA (10), ARX5 (7) and UR5 (3); 20 dual-arm. 18 new tasks plus 12 carried from v1 in revised form; Franka removed; multi-task models only; set-up alignment relaxed.4231
v1 and v2 scores are not comparable · No task is identical across versions: 3 task names recur with changed props, steps or robot (for example three flowers in v1, four in v2). Robots, ranking rules and set-up rules also changed. We found no conversion between the two scales.542+1
Table 30 V2 Simulation (announced) · 2026-08-19: a simulation copy of Table30 v2 in NVIDIA Isaac Lab, with Lightwheel and NVIDIA, aligned in tasks and scoring with the real benchmark. Not released as of 2026-10-10.315
- Last update
- October 202621183
More
Newest Table30 v2 run ended 2026-10-10; leaderboard files regenerated 2026-10-10 07:00 GMT. Last news item 2026-08-19.
- Status
- Active. Table30 v2 has had public runs every month since April 2026. Inferred21
More
Table30 v2 runs every month since April 2026 (141, 298, 41, 162, 71, 289 and 17 public runs from April to October).
Setup
- Runs in
- Real robots1
- Robot
- Two arms, One arm45
More
v2: 20 dual-arm and 10 single-arm tasks. v1: 11 two-arm and 19 single-arm.
- Robot model
- v2: ARX5, UR5, Cobot Magic ALOHA and DOS-W1 (a mobile two-arm robot with Airbot Play arms). v1: UR5 with Robotiq gripper, Franka (Robotiq gripper), Cobot Magic ALOHA and ARX-5. Intel RealSense cameras.1232
More
10 machines at launch (report); 20 by January 2026 (annual report).
- Setting
- Tabletop1
More
'All tasks are executed on the table, or around a table.' One testing site.
- Training data
- 32,980 demonstrations in v24533+1
More
v2: 32,980 demonstration episodes (1,001 to 1,675 per task), about 1.55 TB. v1: up to 1,000 per task; the site lists 25,627 episodes, about 1.57 TB.
The v2 total equals the sum of per-task counts. Reading v1's episode_number field as demonstrations is ours.
- Size
- About 50,000 public test attempts across both versions Inferred3421
More
Our counts. The public run lists hold fewer rollouts than the platform totals, so they appear to be a subset.
v1 platform totals · 41,969 rollouts, 209 test tokens issued and 82 developers submitting by January 2026 (annual report); 'over 80,000' robot task executions by 2026-05-27 (news).203
v1 public run list · 3,985 runs, 39,852 rollouts, 125 user names, 2025-09-20 to 2026-05-25.34
v2 public run list · 1,019 ranked runs, 10,189 rollouts, 51 user names, 2026-04-13 to 2026-10-10.21
Leaderboards · v1 final: 22 entries, all with 30 tasks. v2: 53 entries, 2 with all 30 tasks.1618
- Changes at test
- Only the start positions change. Each test set-up copies a held-out demonstration. Inferred1218
More
Test set-ups copy held-out demonstration episodes. v1 used a pixel-level image overlay; v2 relaxes this to rough alignment and assigns a random robot unit of the right type. Zero-shot object and background tracks are described for v2, but no such scores appear on the public board.
The V2 paper says its competition preview does not score zero-shot or out-of-domain settings and that the out-of-domain mode will not be ranked.
Scoring and access
- Score
- Average success rate over 30 tasks. A progress score breaks ties.12022
More
Per task: success rate over 10 rollouts and a progress score (10 stage points per rollout, minus 0.5 per retry, so 0 to 100). Overall: plain averages over the 30 tasks; ranking by success rate, then progress score.
v1's overall board lists only entries that completed all 30 tasks. v2 averages over all 30 tasks with untested tasks counted as 0, ranks partial entries, and uses the latest ranked run per task (issues.i5). The V2 paper announces rollout timeouts and a 'time to complete' score; no time field appears in the v2 leaderboard data on 2026-10-10.
- Who runs it
- The organisers run the tests Inferred1202+2
More
Step by step: (1) Apply for a test token; download the task data from Hugging Face and fine-tune. (2) Submit a request naming the model and tasks; in v2 a ranked run must cover a whole robot group, and single-task 'test' runs are not ranked. (3) Staff queue the job (hours to days) and tell the user when to have the model running; in v2 the model must declare it is ready before it learns which task comes next. (4) For each rollout a tester places the props to match a reference image from a held-out demonstration and watches the run. (5) The user's program pulls camera images and robot state through RoboChallenge's API and pushes actions into a queue on the robot. (6) Testers score each stage by hand; scores get a second review and can be appealed. (7) Videos and logs are published.
- Error bars
- Not reported Inferred181624
More
Leaderboard data give success rate and score only. PhAIL (2026-05) notes RoboChallenge reports no confidence intervals or paired tests.
- Code licence
- CC-BY-NC-SA-4.01213
More
RoboChallengeInference (client and mock server) LICENSE text is CC BY-NC-SA 4.0; GitHub's API shows NOASSERTION. The organisation's openpi fork is Apache-2.0; the Community repo has no licence. Dexmal's separate OpenDM client is Apache-2.0.
- Data licence
- Unknown333226+1
More
No licence field or text in the Table30 and Table30v2 dataset cards; none of the 63 RoboChallenge datasets on Hugging Face has a licence tag (checked 2026-10-10). The governance document says datasets are to be published under open licences.
- Asset licence
- not applicable Inferred1
More
Physical robots and props; no 3D assets distributed.
- Access
- Testing needs an approved application. The data is open to download.202333
More
Evaluation needs an approved account and test token (209 issued by January 2026). Data downloads are open, without registration.
Sources 34
- 1RoboChallenge technical report, full text v1 (Sections 2-4, Appendix A; Figure 3 tester study)Paper · Oct 2025 · checked 10 Oct 2026
- 2Table30 V2: Evaluating Generalized Models by Real Robots at Scale (CVPR 2026 Workshops, pp. 4461-4467)Paper · Jun 2026 · checked 10 Oct 2026
- 3RoboChallenge news page (dated items 2025-10-17 to 2026-08-19; read from the site's page code)Official site · 19 Aug 2026 · checked 10 Oct 2026
- 4Table 30 V2 (v2.0) benchmark metadata: tasks, robots, prompts, scoring stages, episode countsLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 5Table 30 (v1.0) benchmark metadata: tasks, robots, scoring stagesLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 6DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI (Section 4.2, RoboChallenge results)Paper · Feb 2026 · checked 10 Oct 2026
- 7Qwen-RobotManip Technical Report (v2; Section 6.3.2, Table30-v1 generalist track)Paper · Jun 2026 · checked 10 Oct 2026
- 8StarVLA-alpha: Reducing Complexity in Vision-Language-Action Systems (v2; Section 5, Appendix E)Paper · Apr 2026 · checked 10 Oct 2026
- 9Spirit-v1.5 repository README ('ranks #1 on RoboChallenge Table30' as of 2026-01-11)Repository · Jan 2026 · checked 10 Oct 2026
- 10GigaBrain-0 repository README (news: GigaBrain-0.1 first place on RoboChallenge, 2026-02-09)Repository · Aug 2026 · checked 10 Oct 2026
- 11Dexmal OpenDM README (news 2026-10-05: KDDI Research's VLA-DM0.5 tops Table30 V2; results table)Repository · 5 Oct 2026 · checked 10 Oct 2026
- 12RoboChallengeInference LICENSE (CC BY-NC-SA 4.0) and READMERepository · Oct 2025 · checked 10 Oct 2026
- 13GitHub API: RoboChallenge organisation repositories (stars, licences)Index · 10 Oct 2026 · checked 10 Oct 2026
- 14RoboChallenge CVPR 2026 competition page and task list (same 30 task IDs as v2)Official site · Apr 2026 · checked 10 Oct 2026
- 15RoboChallenge Table 30 V2 simulation page (Lightwheel, NVIDIA Isaac Lab)Official site · 19 Aug 2026 · checked 10 Oct 2026
- 16Table 30 (v1) final leaderboard data (22 entries)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 17Table 30 (v1) per-task leaderboard dataLeaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 18Table 30 V2 leaderboard data (53 entries, per-task results)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 19Dexmal/DM05-Table30v2-W1 model card (DM0.5 checkpoints for Table30 v2)Repository · 6 Aug 2026 · checked 10 Oct 2026
- 202025 RoboChallenge 年度报告 (Annual Report 2025 Q4 - 2026 Q1, Chinese edition)Report · 30 Jan 2026 · checked 10 Oct 2026
- 21Table 30 V2 public run list (1,019 ranked runs with rollout scores)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
- 22RoboChallenge leaderboard page code (ranking order; v2 entries treated as multi-task)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 23RoboChallenge v2 submission form code (ranked versus test runs; whole robot group required)Official site · 10 Oct 2026 · checked 10 Oct 2026
- 24PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology (Related Work)Paper · May 2026 · checked 10 Oct 2026
- 25RoboChallenge Committee Governance ModelOfficial site · Nov 2025 · checked 10 Oct 2026
- 26Hugging Face API: all 63 RoboChallenge datasets (no licence tags)Repository · Apr 2026 · checked 10 Oct 2026
- 27Dexmal official website (北京原力灵机智能科技有限公司; robots and models)Official site · Oct 2026 · checked 10 Oct 2026
- 282025 RoboChallenge Annual Report (English edition)Report · 30 Jan 2026 · checked 10 Oct 2026
- 29RoboChallenge Community README (Chinese edition)Repository · 30 Jan 2026 · checked 10 Oct 2026
- 30RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies (arXiv abstract page, single version v1)Paper · Oct 2025 · checked 10 Oct 2026
- 31QbitAI: authorised reprint of RoboChallenge's Table30 V2 launch announcementSecondary · 24 Mar 2026 · checked 10 Oct 2026
- 32RoboChallenge/Table30v2 dataset card and Hub recordRepository · 5 Jun 2026 · checked 10 Oct 2026
- 33RoboChallenge/Table30 dataset card and Hub recordRepository · 13 Jan 2026 · checked 10 Oct 2026
- 34Table 30 (v1) public run list (3,985 runs with rollout scores)Leaderboard · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
license_data: Table30 and Table30v2 dataset cards, Hugging Face API for all 63 RoboChallenge datasets, governance PDF, annual reports, site code.
cross-site or independent checks: Report, V2 paper, annual reports (Chinese and English), Semantic Scholar list of 34 citing papers (full texts of 10 scanned), web search for critiques in English and Chinese. Found only PhAIL's remark on uncertainty.
zero-shot and time-to-complete scores: v2 leaderboard and run data, benchmark metadata, site code. Not present on 2026-10-10.
official Chinese announcement of v2: Site news, Community README, web search for the WeChat original; only the QbitAI reprint (secondary) was found. The CVPR 2026 Workshops paper is used instead.
statement on v1/v2 comparability: V2 paper, site news, benchmark metadata. No explicit statement; our conclusion is inferred.
Change history
- Created at full depth from primary sources, starting from the checked basic entry. Added the Table30 V2 CVPRW paper, both annual reports, governance model, v1/v2 comparison, leaderboard rules from site code and data, run counts, model reports and tester study. Resolved the basic entry's issue on Qwen-RobotManip: its 'ranks 1st' claim refers to the v1 generalist track, where it is listed as Lira_generalist.
- Published as a full entry.