PAI-Bench

PAI-Bench (Physical AI Bench)

How to read this picture

Tests video generators and video-language models on real-world physical-AI clips: driving, robots, industry, people.1

Sources
Last checked 10 Oct 2026Basic entry21 of 30 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
World models
Robot
No body1
Used by
342
citations
Licence
MIT3

Comparisons with real robots

No comparison found Unknown

Details

Only agreement with human raters was measured: Pearson r = 0.918 between PAI-Bench-G scores and Elo from a pairwise human study (source videos + 8 models), by the benchmark authors. No study links PAI-Bench scores to robot policy success. Looked in the paper, NVIDIA Cosmos-Predict2.5 and Cosmos 3 reports.

Known problems 3

  1. T2v generators span ~4 points on paibench-g vs ~10 on its human eval.

    Claim by NVIDIA Cosmos 3 report (a supporter, not the PAI-Bench authors). Not answered by the PAI-Bench authors as far as we found.4

  2. Per-track counts 1,044 + 600 + 604 + 610 = 2,858, not the stated 2,808. unit of 'case' per

    Our arithmetic; flagged, not resolved. Inferred1

  3. Two models score above the real-video reference row on the generation leaderboard (83.9 an

    Reading of leaderboard data; suggests a ceiling on the automatic metric. Inferred5

Details

About

What it is
Benchmark Inferred6
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
Georgia Tech; CMU (authors Zhou, Huang, Li, Ramanan, Shi). Acknowledgements thank NVIDIA Research, especially the Cosmos team, for support that led to PAI-Bench.1
More

Repo README repeats the same affiliations and acknowledgement.

Released
2025-09 Inferred7
More

GitHub repo created 2025-09-09; HF datasets created 2025-09-12/21; leaderboard Space created 2025-10-10. arXiv v1 is 2025-12-01. Dates from GitHub and HF APIs.

Version
No version tags or numbered releases; datasets last modified 2025-12-10.7
More

GitHub repo has no releases; HF lastModified from API.

Last update
Leaderboard data commit 2026-08-18 'Add Cosmos3-Edge to generation leaderboard'; code commit 2026-06-23 (depth si-RMSE outlier cap; deterministic DOVER scoring).8
More

Code commits at https://github.com/SHI-Labs/physical-ai-bench/commits/main

Status
Active Inferred8
More

Leaderboard updated 2026-08-18; code fixes 2026-06.

Setup

Runs in
Recorded data Inferred1
More

No control loop: generated videos are scored by metrics and a VLM judge; understanding track is multiple-choice QA.

Robot
No body1
Setting
Mixed4
More

Per-domain counts are from NVIDIA's Cosmos 3 report describing PAI-Bench-G; the PAI-Bench paper gives the six domain names.

Size
Unknown61
More

2,808 cases total (headline).6

G: 1,044 video-prompt pairs and 5,636 QA pairs across 6 domains.1

C: 600 videos, 200 clips each from AgiBot (robotics), OpenDV (driving), Ego-Exo4D (egocentric); 1 original + 5 variant captions per video.1

U: 604 QA pairs on 426 videos (physical common sense) + 610 QA pairs from 601 videos (embodied reasoning: RoboVQA, RoboFail, BridgeData, AgiBot, HoloAssist, a proprietary AV dataset).1

Scoring and access

Scored by
Composite index, Automatic judge, Fidelity, Accuracy1
More

0.5/0.5 weighting stated explicitly in Cosmos 3 report (https://arxiv.org/html/2606.02800v4) and Cosmos-Predict2.5 report; consistent with paper tables (e.g. source videos 89.8 domain, 78.0 quality, 83.9 overall).

Who runs it
both: self-reported results, then added to the leaderboard by maintainers via data-file commits Inferred8
Leaderboard
Official leaderboard5
More

Row counts from the Space's JSON data files.

Code licence
MIT3
More

LICENSE file: MIT, Copyright (c) 2025 SHI-Labs.

Data licence
Unknown910
More

PAI-Bench-G: CC-BY-NC-4.09

PAI-Bench-C: MIT; PAI-Bench-U: MIT10

Commercial use
non-commercial for G data; allowed for code, C and U by their stated licences Inferred9
More

Reading of licences; not legal advice.

Published at
CVPR 2026 (README says Oral, news dated 2026-04-09)11
More

CVF open-access page lists the paper in CVPR 2026 proceedings; OpenReview API also lists venue 'CVPR 2026'. Oral status only from README.

Sources 11

  1. 1PAI-Bench: A Comprehensive Benchmark For Physical AI (full text)Paper · Dec 2025 · checked 10 Oct 2026
  2. 2Semantic Scholar API recordIndex · checked 10 Oct 2026
  3. 3SHI-Labs/physical-ai-bench on GitHub (blob)Repository · checked 10 Oct 2026
  4. 4Cosmos 3: Omnimodal World Models for Physical AI (full text)Paper · Jun 2026 · checked 10 Oct 2026
  5. 5shi-labs/physical-ai-bench-leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
  6. 6PAI-Bench: A Comprehensive Benchmark For Physical AIPaper · Dec 2025 · checked 10 Oct 2026
  7. 7SHI-Labs/physical-ai-bench on GitHub (repository)Repository · checked 10 Oct 2026
  8. 8shi-labs/physical-ai-bench-leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
  9. 9shi-labs/physical-ai-bench-generation on Hugging Face (dataset)Repository · checked 10 Oct 2026
  10. 10shi-labs/physical-ai-bench-conditional-generation on Hugging Face (dataset)Repository · checked 10 Oct 2026
  11. 11CVPR 2026 Open Access RepositoryPaper · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).