PAI-Bench
PAI-Bench (Physical AI Bench)
Tests video generators and video-language models on real-world physical-AI clips: driving, robots, industry, people.1
Comparisons with real robots
No comparison found Unknown
Details
Only agreement with human raters was measured: Pearson r = 0.918 between PAI-Bench-G scores and Elo from a pairwise human study (source videos + 8 models), by the benchmark authors. No study links PAI-Bench scores to robot policy success. Looked in the paper, NVIDIA Cosmos-Predict2.5 and Cosmos 3 reports.
Known problems 3
T2v generators span ~4 points on paibench-g vs ~10 on its human eval.
Claim by NVIDIA Cosmos 3 report (a supporter, not the PAI-Bench authors). Not answered by the PAI-Bench authors as far as we found.4
Per-track counts 1,044 + 600 + 604 + 610 = 2,858, not the stated 2,808. unit of 'case' per
Our arithmetic; flagged, not resolved. Inferred1
Two models score above the real-video reference row on the generation leaderboard (83.9 an
Reading of leaderboard data; suggests a ceiling on the automatic metric. Inferred5
Details
About
- What it is
- Benchmark Inferred6
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- Georgia Tech; CMU (authors Zhou, Huang, Li, Ramanan, Shi). Acknowledgements thank NVIDIA Research, especially the Cosmos team, for support that led to PAI-Bench.1
More
Repo README repeats the same affiliations and acknowledgement.
- Released
- 2025-09 Inferred7
More
GitHub repo created 2025-09-09; HF datasets created 2025-09-12/21; leaderboard Space created 2025-10-10. arXiv v1 is 2025-12-01. Dates from GitHub and HF APIs.
- Version
- No version tags or numbered releases; datasets last modified 2025-12-10.7
More
GitHub repo has no releases; HF lastModified from API.
- Last update
- Leaderboard data commit 2026-08-18 'Add Cosmos3-Edge to generation leaderboard'; code commit 2026-06-23 (depth si-RMSE outlier cap; deterministic DOVER scoring).8
More
Code commits at https://github.com/SHI-Labs/physical-ai-bench/commits/main
- Status
- Active Inferred8
More
Leaderboard updated 2026-08-18; code fixes 2026-06.
Setup
- Runs in
- Recorded data Inferred1
More
No control loop: generated videos are scored by metrics and a VLM judge; understanding track is multiple-choice QA.
- Robot
- No body1
- Setting
- Mixed4
More
Per-domain counts are from NVIDIA's Cosmos 3 report describing PAI-Bench-G; the PAI-Bench paper gives the six domain names.
- Size
- Unknown61
More
2,808 cases total (headline).6
G: 1,044 video-prompt pairs and 5,636 QA pairs across 6 domains.1
C: 600 videos, 200 clips each from AgiBot (robotics), OpenDV (driving), Ego-Exo4D (egocentric); 1 original + 5 variant captions per video.1
U: 604 QA pairs on 426 videos (physical common sense) + 610 QA pairs from 601 videos (embodied reasoning: RoboVQA, RoboFail, BridgeData, AgiBot, HoloAssist, a proprietary AV dataset).1
Scoring and access
- Scored by
- Composite index, Automatic judge, Fidelity, Accuracy1
More
0.5/0.5 weighting stated explicitly in Cosmos 3 report (https://arxiv.org/html/2606.02800v4) and Cosmos-Predict2.5 report; consistent with paper tables (e.g. source videos 89.8 domain, 78.0 quality, 83.9 overall).
- Who runs it
- both: self-reported results, then added to the leaderboard by maintainers via data-file commits Inferred8
- Leaderboard
- Official leaderboard5
More
Row counts from the Space's JSON data files.
- Code licence
- MIT3
More
LICENSE file: MIT, Copyright (c) 2025 SHI-Labs.
- Commercial use
- non-commercial for G data; allowed for code, C and U by their stated licences Inferred9
More
Reading of licences; not legal advice.
- Published at
- CVPR 2026 (README says Oral, news dated 2026-04-09)11
More
CVF open-access page lists the paper in CVPR 2026 proceedings; OpenReview API also lists venue 'CVPR 2026'. Oral status only from README.
Sources 11
- 1PAI-Bench: A Comprehensive Benchmark For Physical AI (full text)Paper · Dec 2025 · checked 10 Oct 2026
- 2Semantic Scholar API recordIndex · checked 10 Oct 2026
- 3SHI-Labs/physical-ai-bench on GitHub (blob)Repository · checked 10 Oct 2026
- 4Cosmos 3: Omnimodal World Models for Physical AI (full text)Paper · Jun 2026 · checked 10 Oct 2026
- 5shi-labs/physical-ai-bench-leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
- 6PAI-Bench: A Comprehensive Benchmark For Physical AIPaper · Dec 2025 · checked 10 Oct 2026
- 7SHI-Labs/physical-ai-bench on GitHub (repository)Repository · checked 10 Oct 2026
- 8shi-labs/physical-ai-bench-leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
- 9shi-labs/physical-ai-bench-generation on Hugging Face (dataset)Repository · checked 10 Oct 2026
- 10shi-labs/physical-ai-bench-conditional-generation on Hugging Face (dataset)Repository · checked 10 Oct 2026
- 11CVPR 2026 Open Access RepositoryPaper · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).