RoboWorld

How to read this picture

Runs robot policies inside a learned video world model trained on DROID; a VLM judge scores task progress.1

Sources
Last checked 10 Oct 2026Basic entry15 of 22 facts checked at the sourceNext check 8 Apr 2027
Runs in
AI simulator1
Checked against real robots
Checked
Skill
Handling objects
Robot
One arm1
Used by
32
citations

Comparisons with real robots

Checked1

Details

8 policies; Pearson r 0.989 and Spearman rho 0.970 between RoboWorld scores (GPT-4o judge, 0-5 progress rubric) and the RoboArena real-world leaderboard (snapshot 2026-02-26). Binary-success scoring gives rho 0.922. Synthetic 'extreme' environments still r 0.970 vs RoboArena. Measured by the RoboWorld authors against real data collected by a different group (RoboArena); no third-party replication found.

Known problems 1

  1. Stated limitations

    Long-horizon, contact-rich manipulation with object consistency remains hard for video world models. Correlation rests on 8 policies and one VLM judge (GPT-4o).1

Details

About

What it is
Simulator Inferred3
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
KAIST; Config1
More

Affiliations in arXiv v4 and project page (1 KAIST, 2 Config). What kind of organisation 'Config' is was not checked.

Released
2026-07 (arXiv v1 2026-07-01)3
Version
arXiv v4; code 'coming soon'4
More

Project page shows 'Code coming soon'; no repo linked from paper or page.

Last update
arXiv v4 2026-07-15 (v2 07-13, v3 07-14)3
More

What changed between versions is not stated (file sizes nearly identical).

Status
Active Inferred3
More

Paper revised four times in July 2026; code promised.

Setup

Runs in
AI simulator1
More

Closed-loop: policy acts on generated frames; world model predicts next frames from actions.

Robot
One arm1
Robot model
DROID setup (Franka Panda; two external views + one wrist view, tiled 2x2)1
More

Conditioned on end-effector Cartesian position; adapter for joint-velocity policies.

Setting
Mixed Inferred1
More

Initial frames come from RoboArena episodes (many sites); plus 8 synthetic environments made with an image editor.

Size
Unknown1
More

8 policies; 4,186 rollouts; 30 s per rollout; about 100 H100 GPU-hours for the full RoboArena replication1

Synthetic: 175 initial observations -> 746 valid initial conditions across 8 environments1

Scoring and access

Scored by
Automatic judge, Progress score1
More

GPT-4o judge with a 0-5 task-progress rubric scored from fixed external views (wrist view treated as less reliable).

Leaderboard
None4
More

No leaderboard on page or paper.

Code licence
Unknown
More

Code not released ('Code coming soon'); looked at project page and arXiv links.

Data licence
Unknown
More

No new dataset released; uses DROID (training) and RoboArena data dump (MIT, see RoboArena record). Paper text is CC BY 4.0.

Published at
ICML 2026 F2S Workshop on Long-Horizon Video Generation4
More

Stated on project page; arXiv comments give only the project link.

Sources 4

  1. 1RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation (full text)Paper · Jul 2026 · checked 10 Oct 2026
  2. 2Semantic Scholar API recordIndex · checked 10 Oct 2026
  3. 3RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy EvaluationPaper · Jul 2026 · checked 10 Oct 2026
  4. 4RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy EvaluationOfficial site · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).