EWMBench

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

How to read this picture

EWMBench is a benchmark for world models, which here are AI models that generate videos of a robot doing a task. It scores how closely those videos match real recordings of the same task.123

Sources
Last checked 10 Oct 2026Full entry47 of 62 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
World models
Robot
The videos show a two-armed robot.14
Used by
7 papers or contests by October 2026567+5
41 citations
Licence
CC-BY-NC-SA-4.0213
Commercial use: not allowed

What a score here does not tell you Inferred

  1. Whether the world model helps a robot do its task.EWMBench scores how much the videos look like the recordings. It does not measure robot success.
  2. How a model does on the full test set.Only 21 of the 100 test episodes are public.
  3. Whether the automatic judges are reliable.Only one small human check has been done, and it reports no statistic.
ChartPublished scores over time
TOP 5% OF THE SCALE22.533.544.555.566.577.582026EnerVerse_FT · 4.701 · 2025-05LTX_FT · 4.5493 · 2025-05Kling-1.6 · 3.8698 · 2025-05Hailuo I2V-01-live · 3.4125 · 2025-05COSMOS-7B · 3.2872 · 2025-05OpenSora 2.0 · 3.1392 · 2025-05LTX-Video · 2.9676 · 2025-05Qwen-RobotWorld · 4.6 · 2026-06LVP · 4.05 · 2026-06Sora2 · 3.89 · 2026-06Kling 2.6 · 3.85 · 2026-06Veo3 · 3.49 · 2026-06EnerVerse_FT 4.701EnerVerse_FT 4.701
Score: Overall score, the sum of 8 metrics, out of 8. Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.17

Comparisons with real robots

StudyResultWhat was comparedDone by
Human ranking study (EWMBench paper)
May 2025
The order was the same for all 4 models. No statistic was reported.11415This was not a real-robot comparison. People ranked videos from 4 video models, and their combined order was compared with the orders from EWMBench and VBench. The study’s authors described the result as “align more closely with human judgments than VBench”.The benchmark’s authors

Benchmarks built on EWMBench: AgiBot World Challenge, World Model track.89

Our assessment Opinion

EWMBench measures how much the videos look like real recordings. It does not measure robot success.

Reasoning

EWMBench shows how closely a model's video of a robot task matches a recording, as judged by automatic tools. It does not show whether the model would help a robot succeed, and no study has linked its scores to robot results.

Confidence: high

Published numbers are hard to compare or reproduce.

Reasoning

Treat published EWMBench numbers as rough guides. The full test set is not public, later papers use a 21-sample subset or adapted metrics, the released code differs from the paper, and the builder's own papers attach the same numbers to different model names.

Confidence: high

The human check is too small to show that the automatic judges are reliable.

Reasoning

The human check supports the ranking of four models. It is too small to show that the automatic judges are reliable across models, and its figure values differ from the main table.

Confidence: medium

It is used mostly through AgiBot's contests. Its licence is non-commercial.

Reasoning

Its metrics are still used, mainly to score AgiBot's world-model contests, which use three of the eight. The non-commercial licence limits company use outside research.

Confidence: medium

Known problems 6

  1. Most of the test set is not public

    Only 21 of the paper's 100 test episodes are public.16177+1

    Details

    The paper scores models on 100 episodes from 10 tasks. The public ground truth holds 21 samples from 7 task categories. A repository collaborator called it a verification subset on 2025-06-20 and said the full version would follow. Users asked again in July, October and November 2025 without a reply, and the Hugging Face data has not changed since 2025-05-16. Later papers such as Qwen-RobotWorld score models on the 21-sample set, so their numbers are not on the same footing as the paper's.

  2. The same scores appear under different model names

    A later AgiBot paper lists EnerVerse_FT's exact scores under another model's name.1618+4

    Details

    EWMBench Table 2 credits EnerVerse_FT with Scene 0.9427, Motion 1.6676, Semantics 2.0907 and Overall 4.7010. Genie Envisioner, a later AgiBot paper whose authors include all eight EWMBench authors, lists the same four numbers for GE-Base (Figure 17, unchanged from its v1 of 2025-08 to v3 of 2025-11). In EWMBench's human study the top model is LTX_FT; Genie Envisioner shows the same EWMBench and VBench bars with GE-Base in that place. Neither paper explains this. Genie Envisioner says GE-Base is built on LTX-Video. Separately, the Qwen-RobotWorld report evaluates on the 21-sample set, yet its Cosmos row is identical to the paper's COSMOS row (0.7963 to 0.7333 on all eight metrics), which came from the 100-episode protocol.

  3. It does not test whether a world model helps a robot

    Other research groups note that it does not test whether a world model helps a robot decide or act.212223+1

    Details

    Later benchmark papers say EWMBench measures the generated video and does not test whether a world model helps a robot decide or act. WorldArena 2.0 (2026-05) says such benchmarks 'do not assess whether generated dynamics support embodied decision-making'. Wow, wo, val! (2026-01) says EWMBench does not assess planning and execution. WorldArena (2026-02) marks it as lacking data-engine, policy-evaluation and action-planner tests. WorldArena also found that its own video-quality score correlated with downstream robot tasks at only r 0.600 and r 0.360, which suggests video scores and usefulness can diverge; that finding is about WorldArena's score, not EWMBench's.

  4. The human check is small and gives no statistic

    The check against human rankings covers only four models. The paper does not say how many people took part and gives no statistic.11514

    Details

    Annotators ranked videos from four models (LTX_FT, Kling-1.6, Hailuo I2V-01-live, OpenSora-2.0), giving 3, 2 and 0 points to the best, second-best and worst. The paper gives no number of annotators or videos and no agreement statistic. The aggregated human order matched EWMBench's order for all four models, while VBench put Hailuo first. The EWMBench scores in the same figure (5.49, 4.75, 4.30, 4.03) differ from Table 2's Overall scores for these models (4.5493, 3.8698, 3.4125, 3.1392), and the paper does not explain the difference. The automatic judges (Qwen2.5-VL-7B for captions and logic) are not checked against human labels; only the gripper detector is tested (recall 0.91667, precision 1.0 on held-out frames).

  5. The released code differs from the paper's method

    The public scoring code differs from the method the paper describes. Inferred252627+4

    Details

    In our reading of the released toolkit, the summary script averages over all generated videos and has no best-of-3 selection. BLEU and CLIP are both computed on the one-line video summary, while the paper describes CLIP on step-by-step descriptions. Motion scores are divided by fixed constants that the paper does not report. A variable-name typo (trail_id for trial_id) attaches each Logics value to the wrong video or drops it from the summary table. Genie Envisioner also describes the ground-truth paths as manually annotated, while the EWMBench toolkit detects them automatically. Scores from the public code may therefore differ from Table 2.

  6. A still video can score well on scene stability

    The paper itself notes that videos with no motion can get high scene scores.1

    Details

    The paper notes that visually plausible but static videos may score high on scene consistency while lacking meaningful motion. The Overall score adds the scene score to the motion and semantic scores, so part of the total does not depend on doing the task. OpenSora and LTX, which the paper says often produce static videos, scored 0.9210 and 0.9156 on scene consistency, above Kling (0.8888).

Details

About

What it is
Benchmark1
More

The paper presents a curated dataset, fixed metrics and an open evaluation toolkit, and calls it a benchmark.

Built by
AgiBot, Shanghai Jiao Tong University, Harbin Institute of Technology, National University of Singapore / CUHK MMLab12915
More

Authors: Yue Hu, Siyuan Huang, Yue Liao, Shengcong Chen, Pengfei Zhou, Liliang Chen, Maoqing Yao, Guanghui Ren. arXiv v1 prints the first name as 'Hu Yue'.

AgiBot · Six of eight authors list AgiBot. The BMVC PDF says the three authors with other affiliations did the work while employed at AgiBot.2915

Shanghai Jiao Tong University · Affiliation of Siyuan Huang.129

Harbin Institute of Technology (Weihai) · Second affiliation of Yue Hu.15

MMLab-CUHK or National University of Singapore · Affiliation of Yue Liao. CONFLICT: arXiv says MMLab-CUHK; the BMVC page and PDF say National University of Singapore.12915

Released
May 2025, at BMVC 2025301317+1
More

arXiv v1 on 2025-05-14. GitHub repository created 2025-05-14; Hugging Face data created 2025-05-15. Published at BMVC 2025 (Sheffield, 24 to 27 November 2025).

The BMVC page header says '35th' conference while its citation block says '36th'.

Version
No releases. The paper is at version 2.301331+1
More

No tags or releases. Paper arXiv v2 (2025-05-18). Evaluation model weights are labelled v0.1.

We compared the arXiv v1 and v2 texts: they differ in author-name order and citation style, and every decimal number is the same. The BMVC camera-ready has the same Table 2.

Last update
June 2025. The code has not changed since then.3317
More

Last code commit 2025-06-13. The Hugging Face data has not changed since 2025-05-16.

The commit 'add psnr and ssim' is dated 2025-06-11. The full ground-truth set promised in issue #1 has not appeared (see issues.i1).

Status
No code or data changes since June 2025. Its metrics are still reused. Inferred331634+1
More

No code or data change since June 2025. AgiBot still uses three of its metrics in its 2026 contest.

Last commit 2025-06-13. Users asked about the missing full test set in issue #1 on 2025-07-09, 2025-10-16 and 2025-11-04 without a reply; issue #2 (2025-11-10) is unanswered. The basic entry said 'maintained'; changed because nothing has been updated for over a year.

Setup

Runs in
Recorded data1
More

Generated videos are compared with recorded real-robot videos. No robot moves and no policy runs.

Robot
The videos show a two-armed robot. Inferred14
More

Videos show a two-armed robot from its head camera. The model under test is a video generator.

The EWMBench paper does not name the robot; its trajectory metrics track left and right end effectors. The source data, AgiBot World, was collected with 'dual-arm humanoid robots' on mobile bases (AgiBot World paper).

Setting
Kitchen, Whole home, Retail or logistics, Industrial Inferred1
More

Mapped by us from the task list: toaster, pouring water, cutlery, microwave (kitchen); showerhead, drawer, bottle cleaning, ice (home); freezer restocking (retail); detergent packing (industrial).

Tasks
10 tasks1
More

10 manipulation tasks from AgiBot World, each split into 4 to 10 captioned sub-actions

Tasks (Appendix A.2.1): retrieving toast from a toaster, pouring water, setting cutlery, restocking a freezer, producing ice, packing laundry detergent, cleaning bottles, heating food in a microwave, installing a showerhead, storing objects in a drawer.

Changes at test
Not described in the paper Inferred16
More

The paper does not describe held-out conditions. Test episodes were chosen to have varied arm paths.

Episodes were picked by a greedy rule that maximises differences between voxelised gripper paths. Genie Envisioner says the 10 tasks were left out of GE-Base pre-training; the EWMBench paper does not say whether other tested models saw these tasks. The two top models are described as fine-tuned for embodied scenes, without naming the data.

Scoring and access

Scored by
Fidelity, Automatic judge, Composite index1
More

Fidelity: comparison with the recorded video and gripper path. Automatic judge: Qwen2.5-VL captions and logic check. Composite: the Overall sum.

Score
The sum of 8 automatic scores, with a maximum of 8 Inferred12627+3
More

Scene (1 metric): similarity of fine-tuned DINOv2 features between frames. Motion (3): Hausdorff distance, normalised dynamic time warping (nDTW) and a velocity and acceleration match, all on gripper paths found by a fine-tuned YOLO-World detector. Semantics (4): BLEU and CLIP similarity of captions written by Qwen2.5-VL-7B, a yes-or-no logic check by the same model, and diversity across generations. Each metric is scaled to 0 to 1. The Overall score is their sum, so the maximum is 8.

Metric definitions are verified in the paper and code. That Overall is a sum is our arithmetic: the 'Avg.' columns in Table 2 are sums (EnerVerse_FT 0.9427 + 1.6676 + 2.0907 = 4.7010). In the released code the motion scores are divided by fixed constants (24.979, 22.519, 50.202) and capped at 1; the paper does not give these constants. Gripper paths are 2D image positions. The logic prompt tells the judge to count any human hand in the video as a violation.

Trials
Best of 3 videos per episode1
More

Each model makes 3 videos per episode. Only the video whose gripper path is closest to the recording (by Hausdorff distance) is scored. 2,100 videos for 7 models.

Released code · The summary script averages over all generated videos. We found no best-of-3 selection step.2526

Logics values fit 90 videos · Every Logics value in Table 2 is a whole multiple of 1/90 (for example 0.9778 = 88/90). That fits 30 samples times 3 generations and does not fit 100 or 300 videos.1

Genie Envisioner · Three samples per instruction; the one with the lowest Hausdorff distance is kept.6

Qwen-RobotWorld · Evaluates on the public set of 21 samples across 7 tasks.7

Who runs it
Each team tests its own model Inferred2
More

Users run the toolkit on their own generations. AgiBot's contest tracks run an organiser test server, but only with three of the metrics.

Error bars
Not reported Inferred167
More

Table 2 gives single numbers. Genie Envisioner and Qwen-RobotWorld also give single numbers without error bars.

Leaderboard
None. Scores are only in papers. Inferred238
More

No leaderboard in the repository, dataset card or paper. AgiBot's challenge leaderboards use three EWMBench metrics on their own test data (see facts.derived_benchmarks).

Code licence
CC-BY-NC-SA-4.0213
More

README: all data and code in the repository are under CC BY-NC-SA 4.0. There is no LICENSE file (HTTP 404) and GitHub detects no licence.

Data licence
CC-BY-NC-SA-4.0337
More

Hugging Face dataset card metadata; not gated. The ground-truth clips come from AgiBot World, whose README also states CC BY-NC-SA 4.0.

Asset licence
The evaluation model weights (fine-tuned DINOv2 and YOLO-World) are CC BY-NC-SA 4.0 on Hugging Face.31136+1
More

The detector is fine-tuned from Ultralytics yolov8s-worldv2 (paper A.1), and the toolkit imports the ultralytics package, whose LICENSE file is AGPL-3.0. The project does not say how the two licences combine. Not legal advice.

Access
Open, but only a subset of the test set is public.2331
More

Code on GitHub; data and weights on Hugging Face without a gate or registration. Only the 21-sample subset of the ground truth is public (issues.i1).

Commercial use
Not allowed Inferred2331
More

Code, data and weights all carry the NC (non-commercial) clause of CC BY-NC-SA 4.0. Not legal advice.

Sources 38

  1. 1EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models (full text, arXiv v2)Paper · May 2025 · checked 10 Oct 2026
  2. 2EWMBench GitHub READMERepository · May 2025 · checked 10 Oct 2026
  3. 3agibot-world/EWMBench dataset card and file listing (Hugging Face)Repository · May 2025 · checked 10 Oct 2026
  4. 4AgiBot World Colosseo paper (robots used for data collection)Paper · Mar 2025 · checked 10 Oct 2026
  5. 5Semantic Scholar API record for arXiv:2505.09694Index · 10 Oct 2026 · checked 10 Oct 2026
  6. 6Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation (v3; EWMBench sections)Paper · Aug 2025 · checked 10 Oct 2026
  7. 7Qwen-RobotWorld Technical Report (Section 5.1.1 and Table 2: EWMBench)Paper · Jun 2026 · checked 10 Oct 2026
  8. 8AgiBot World Challenge @ ICRA 2026, World Model track test server (evaluation rules)Leaderboard · Feb 2026 · checked 10 Oct 2026
  9. 9World Model baseline README for AgiBot World Challenge @ IROS 2025 (commit 62eb4ebfec)Repository · 15 Jul 2025 · checked 10 Oct 2026
  10. 10Pelican-Sim 1.0: A General World Model Simulator for Embodied IntelligencePaper · Sep 2026 · checked 10 Oct 2026
  11. 11EVEWorld: Physical Evolution Supervision for Embodied World ModelsPaper · Oct 2026 · checked 10 Oct 2026
  12. 12ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World ModelsPaper · Aug 2026 · checked 10 Oct 2026
  13. 13GitHub API: AgibotTech/EWMBench (stars, forks, created, pushed, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
  14. 14EWMBench Figure 6 (human rank and EWMBench vs VBench scores)Paper · May 2025 · checked 10 Oct 2026
  15. 15EWMBench, BMVC 2025 camera-ready PDFPaper · Nov 2025 · checked 10 Oct 2026
  16. 16EWMBench GitHub issue #1: Missing samples (collaborator reply 2025-06-20)Repository · Jun 2025 · checked 10 Oct 2026
  17. 17Hugging Face Hub API record for agibot-world/EWMBench (downloads, likes, created, last modified, commits)Index · 10 Oct 2026 · checked 10 Oct 2026
  18. 18Genie Envisioner Figure 17 (EWMBench results table naming GE-Base)Paper · Aug 2025 · checked 10 Oct 2026
  19. 19Genie Envisioner Figure 18 (human, VBench and EWMBench ranks with GE-Base)Paper · Aug 2025 · checked 10 Oct 2026
  20. 20Genie Envisioner arXiv abstract page (authors, v1 2025-08-07 to v3 2025-11-04)Paper · Aug 2025 · checked 10 Oct 2026
  21. 21WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform (Section 2.2)Paper · May 2026 · checked 10 Oct 2026
  22. 22Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing TestPaper · Jan 2026 · checked 10 Oct 2026
  23. 23WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsPaper · Feb 2026 · checked 10 Oct 2026
  24. 24How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric PositionPaper · Jun 2026 · checked 10 Oct 2026
  25. 25EWMBench scoring code: EWMBench/__init__.py (result merging and means)Repository · 13 Jun 2025 · checked 10 Oct 2026
  26. 26EWMBench scoring code: trajectory_consistency.py (HSD, nDTW, DYN, normalisation constants)Repository · Jun 2025 · checked 10 Oct 2026
  27. 27EWMBench scoring code: caption.py (Qwen2.5-VL prompt and logic check)Repository · Jun 2025 · checked 10 Oct 2026
  28. 28EWMBench scoring code: semantics.py (BLEU and CLIP on captions)Repository · Jun 2025 · checked 10 Oct 2026
  29. 29EWMBench, BMVC 2025 proceedings pagePaper · Nov 2025 · checked 10 Oct 2026
  30. 30EWMBench arXiv abstract page (submission history)Paper · May 2025 · checked 10 Oct 2026
  31. 31agibot-world/EWMBench-model (Hugging Face model card: fine-tuned DINOv2 and YOLO-World weights)Repository · May 2025 · checked 10 Oct 2026
  32. 32EWMBench full text, arXiv v1 (for version comparison)Paper · May 2025 · checked 10 Oct 2026
  33. 33EWMBench commit historyRepository · 13 Jun 2025 · checked 10 Oct 2026
  34. 34EWMBench GitHub issues listRepository · Nov 2025 · checked 10 Oct 2026
  35. 35EWMBench scoring code: scene_consistency.py (DINOv2 frame similarity)Repository · Jun 2025 · checked 10 Oct 2026
  36. 36EWMBench preprocessing: processing/detection_tracking.py (ultralytics YOLO, 2D gripper paths)Repository · Jun 2025 · checked 10 Oct 2026
  37. 37AgiBot World README (licence section)Repository · 2025 · checked 10 Oct 2026
  38. 38Ultralytics LICENSE (AGPL-3.0)Repository · 2026 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: EWMBench arXiv v1 and v2 and BMVC camera-ready (full text); GitHub README and issues; Hugging Face cards; Genie Envisioner v3 (2508.05635); Qwen-RobotWorld (2606.17030); WorldArena (2602.08971); WorldArena 2.0 (2605.17912); Wow, wo, val! (2601.04137); decision-making position paper (2606.15032); Pelican-Sim 1.0 (2609.12036); all 41 Semantic Scholar citing papers scanned by title, 21 opened and searched for 'EWMBench'; web searches for EWMBench critiques and correlation with policy success. No paired comparison with robot results found.

validity (agreement with humans): Paper Section 4.2 and Figure 6 (arXiv v2 and BMVC PDF, figure read as an image); Genie Envisioner Figure 18. No annotator count, video count or agreement statistic in either.

test_set (full ground truth): Hugging Face dataset tree and commit list (last change 2025-05-16), datasets-server row keys, GitHub issue #1 and #2, README. Only the 21-sample subset is available.

leaderboard: Paper, GitHub README, Hugging Face dataset and model cards, AgiBot ICRA 2026 world-model test server page.

license_code: GitHub contents API for LICENSE (404), repository licence field (null), README licence section.

issues.i2 (explanation of GE-Base label): Genie Envisioner v1 and v3 text and figures (figure files are byte-identical across versions); EWMBench paper. Neither explains the relabelling.

Change history

  1. Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/world-models.json. Changes from the basic entry: status set to dormant (no update for over a year); the '4,644 rows' are frames, so test-set size now separates the paper's 100 episodes from the public 21; embodiment adds humanoid and scene adds kitchen; alias WMBM added. New findings: identical scores credited to EnerVerse_FT and GE-Base, copied Cosmos row in Qwen-RobotWorld, code-versus-paper differences, figure-versus-table mismatch in the human study, and the AGPL-3.0 upstream of the detector.
  2. Published as a full entry.