MVPBench

MVPBench (Minimal Video Pairs)

How to read this picture

Video question set on physical understanding; each question comes as a near-identical video pair with opposite answers.1

Sources
Last checked 10 Oct 2026Basic entry16 of 21 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
Reasoning
Robot
No body2
Used by
303
citations
Licence
Unclear4

Comparisons with real robots

No comparison found Unknown

Details

Not built to predict robot performance. Searched paper, repo, HF card.

Known problems 2

  1. The benchmark is designed against shortcut solutions; authors show language-only, single-f

    Design feature, recorded for context.1

  2. Mini split described as '9k examples' (readme, hf card) and '~5k items' (leaderboard text)

    Likely different units (rows vs pairs); not resolved. Inferred5

Details

About

What it is
Benchmark Inferred6
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
FAIR at Meta (all authors); first author also Mila and McGill University, work done during internship1
Released
2025-06 (arXiv v1 2025-06-11)6
More

HF dataset object created 2025-03-24 (API), announced with the blog 2025-06-11.

Version
arXiv v1 only; TMLR 2025 publication6
Last update
OpenReview: 'Accepted by TMLR', publication date 2025-12-01. Last repo commit 2025-09-22 (download script fix).7
More

OpenReview API search; Semantic Scholar venue also 'Trans. Mach. Learn. Res.'; commit date from GitHub API.

Status
Maintained Inferred2

Setup

Runs in
Recorded data Inferred1
Robot
No body2
Setting
Mixed2
Size
Unknown68
More

55K examples from 9 sources; human 92.9%, best open-source model 40.2%, random 25%.6

HF card: full 54,828 rows; mini 18,416 rows (5,724 + 8,680 + 2,000 + 2,012).8

Scoring and access

Scored by
Accuracy2
Leaderboard
Official leaderboard9
Code licence
CC-BY-NC-4.0 (repo LICENSE file)4
More

GitHub reports 'Other'; README says benchmark released under this LICENSE.

Data licence
HF card metadata: Apache-2.0. Videos are not hosted 'for legal reasons'; users download them from the nine original sources and accept their licences.8
More

Conflicts with repo LICENSE (CC-BY-NC-4.0). Flagged.

Sources 9

  1. 1A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs (full text)Paper · Jun 2025 · checked 10 Oct 2026
  2. 2facebookresearch/minimal_video_pairs on GitHub (repository)Repository · checked 10 Oct 2026
  3. 3Semantic Scholar API recordIndex · checked 10 Oct 2026
  4. 4facebookresearch/minimal_video_pairs on GitHub (blob)Repository · checked 10 Oct 2026
  5. 5facebook/physical_reasoning_leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026
  6. 6A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video PairsPaper · Jun 2025 · checked 10 Oct 2026
  7. 7https://api2.openreview.net/notes/search?term=Minimal%20Video%20Pairs%20shortcut-awareIndex · checked 10 Oct 2026
  8. 8facebook/minimal_video_pairs on Hugging Face (dataset)Repository · checked 10 Oct 2026
  9. 9facebook/physical_reasoning_leaderboard on Hugging Face (space)Leaderboard · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).