RBench

How to read this picture

Scores video generators on robot task videos across five task types and four robot body types.1

Sources
Last checked 10 Oct 2026Basic entry15 of 22 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
World models
Robot
One arm, Two arms, Humanoid, Legged2
Used by
333
citations

Comparisons with real robots

No comparison found Unknown

Details

No link to robot execution. Human alignment only: Spearman rho = 0.96 (p < 1e-3) between RBench and pairwise human scores on 10 models, 30 participants, by the authors. README lists 'Embodied Execution Evaluation' via inverse dynamics model as a to-do, not done.

Details

About

What it is
Benchmark Inferred4
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
Peking University; ByteDance Seed (author affiliations 1 and 2)1
Released
2026-01 (arXiv v1 2026-01-21; repo created 2026-01-21; HF RBench dataset created 2026-01-15)4
More

Repo/HF dates from APIs.

Version
arXiv v1 only; published version in ICML 2026 (PMLR 306:24240-24283). No release tags.5
Last update
Leaderboard data updated 2026-09-08 (commit 'Update leaderboard.json'); repo commit 2026-06-01 'Add Cosmos 3 RBench news'; RoVid-X released 2026-05-21.6
More

Dates from HF Space commit API and GitHub API.

Status
Active Inferred6
More

Leaderboard updated 2026-09-08.

Setup

Runs in
Recorded data Inferred1
More

Generated videos scored by MLLM judges and vision operators; no control loop.

Robot
One arm, Two arms, Humanoid, Legged2
Setting
Mixed Inferred1
Size
650 image-text pairs = 250 task-oriented (50 per category) + 400 embodiment-specific (100 per type).2
More

Same in paper Sec. 3.1.

Scoring and access

Scored by
Automatic judge, Composite index1
Leaderboard
Official leaderboard7
Code licence
Unknown
More

No LICENSE file (404) and GitHub reports no licence. README badge says 'Apache-2.0' but links to placeholder 'YOUR_LINK'.

Data licence
CC-BY-4.0 (RBench evaluation set)2
More

HF dataset card metadata; not gated. Reference images come from public datasets and online sources; upstream terms not restated.

Published at
ICML 2026, PMLR vol. 306, pp. 24240-24283; also ICLR 2026 Workshop on World Models (OpenReview).5
More

Workshop venue from OpenReview API search (https://api2.openreview.net/notes/search).

Sources 7

  1. 1Rethinking Video Generation Model for the Embodied World (full text)Paper · Jan 2026 · checked 10 Oct 2026
  2. 2DAGroup-PKU/RBench on Hugging Face (dataset)Repository · checked 10 Oct 2026
  3. 3Semantic Scholar API recordIndex · checked 10 Oct 2026
  4. 4Rethinking Video Generation Model for the Embodied WorldPaper · Jan 2026 · checked 10 Oct 2026
  5. 5Rethinking Video Generation Model for the Embodied WorldPaper · checked 10 Oct 2026
  6. 6DAGroup-PKU/RBench-Leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
  7. 7DAGroup-PKU/RBench-Leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).