Cosmos-HumanEval

Cosmos-HumanEval (Cosmos-HUE)

How to read this picture

NVIDIA human-rating protocol: annotators answer yes/no physics and fidelity questions about generated videos.1

Sources
Last checked 10 Oct 2026Basic entry13 of 19 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
World models
Robot
No body2

Comparisons with real robots

No comparison found Unknown

Details

No evidence that HUE scores predict robot policy success. The report validates HUE by real-video ground truth (93.6 T2V, 94.4 I2V) and binomial confidence intervals, not by robot outcomes. Looked in Cosmos 3 report and dataset card.

Details

About

What it is
Benchmark Inferred3
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
NVIDIA Corporation (dataset owner); introduced in the Cosmos 3 report by NVIDIA.2
Released
2026-052
More

HF dataset created 2026-05-05 (API); card lists 'Dataset Creation Date: 2026-05-20'. Cosmos 3 report v1 2026-06-01.

Version
v1.2-opensource ('HUE-PaiBench v1.2'); the card calls it the publicly releasable subset of NVIDIA's HUE question bank.2
Last update
HF dataset lastModified 2026-06-09; Cosmos 3 report v4 submitted 2026-06-23.3
More

HF date from API.

Status
Active Inferred1
More

Report says question-bank refinement is ongoing.

Setup

Runs in
Recorded data Inferred1
More

Humans rate generated videos; no control loop.

Robot
No body2
Setting
Mixed2
Size
Unknown21
More

197 samples (100 I2V, 97 T2V); 2,957 questions; per category: Physical Laws 822, Visual Integrity 771, Semantic Alignment 754, Geometric Reasoning 610.2

T2V evaluation pool: 100 prompts sampled from PAIBench-G, 5 seeds each, up to 20 questions per video, up to 10,000 binary observations per checkpoint.1

Scoring and access

Scored by
Human rating1
Who runs it
The organisers run the tests2
Leaderboard
None. Scores are only in papers.1
Code licence
Unknown
More

Dataset holds only two JSON question files; no scoring code found on the card. Cosmos 3 abstract says code, checkpoints and an evaluation benchmark are released under OpenMDW-1.1, but we did not find HUE scoring code.

Data licence
custom: OpenMDW-1.1; card adds that the data was created in part with GPT-5.2 and may not be used to develop or train AI/ML systems. Card also says it is ready for commercial or non-commercial use.2

Sources 3

  1. 1Cosmos 3: Omnimodal World Models for Physical AI (full text)Paper · Jun 2026 · checked 10 Oct 2026
  2. 2nvidia/Cosmos-HumanEval-v1 on Hugging Face (dataset)Repository · checked 10 Oct 2026
  3. 3Cosmos 3: Omnimodal World Models for Physical AIPaper · Jun 2026 · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).