VSI-Bench

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

How to read this picture

VSI-Bench is a set of 5,130 questions about videos of real rooms, used to test how well multimodal AI models (models that take in video and text) understand space. Models only answer questions, and no robot moves.12

Sources
Last checked 10 Oct 2026Full entry61 of 72 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
Reasoning
Robot
No body1
Used by
11 reports345+8
814 citations
Licence
Apache-2.01415
Commercial use: not allowed

What a score here does not tell you Inferred

  1. Whether a robot using the model will succeed at tasks.No study links VSI-Bench scores to robot results.
  2. Whether the model actually used the video.A model trained on the questions alone, without video, gained 17.9 points.
  3. How a score compares with scores in other papers.Papers use different numbers of video frames and different ways of averaging.
ChartPublished scores over time
TOP 5% OF THE SCALE3040506070809010020252026Gemini-1.5 Pro · 45.4 · 2024-12InternVL3.5-241B-A28B · 69.5 · 2025-08GPT-5 (run by Google) · 52.9 · 2025-10Qwen3-VL-235B-A22B · 60 · 2025-11Cambrian-S-7B · 67.5 · 2025-11SenseNova-SI (8B) · 68.8 · 2025-11SpaceMind · 69.6 · 2025-11Gemini 3 Pro · 56 · 2026-03HY-Embodied-0.5 MoE-A32B · 68.3 · 2026-04SpaceMind++ · 73.2 · 2026-05Gemini-1.5 Pro 45.4SpaceMind++ 73.2
Score: Average of the 8 task scores. Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.153+9

Comparisons with real robots

Not checked Inferred1918

Details

MV-RoboBench (ICLR 2026) reports that strong scores on general single-view spatial benchmarks do not reliably carry over to robotic spatial questions in its own set; that compares question sets, not robot runs. Vlaser (ICLR 2026) found that gains on embodied-reasoning benchmarks, VSI-Bench among them, did not carry over to closed-loop robot control in simulation.

Our assessment Opinion

VSI-Bench measures how well a model answers spatial questions from video. No study links its scores to robot success.

Reasoning

VSI-Bench measures how well a model answers questions about the layout of real rooms from video. No study links its scores to robot task success, and in one study (Vlaser) gains on such benchmarks did not carry over to robot control in simulation.

Confidence: medium

Check the frame count, the averaging and the training data before comparing scores.

Reasoning

Before comparing two VSI-Bench numbers, check the frame count, the averaging method and whether the model was trained on same-template data. Each of these has moved scores by between 5 and 18 points.

Confidence: high

Humans score higher than models on layout questions, but several models beat them on size and distance.

Reasoning

The human score of 79.2 is an uneven reference. Humans scored 94 to 100 on layout and order tasks but 45.9 to 60.4 on size and distance estimates, where several 2026 models already score higher. It also rests on 50 questions per task.

Confidence: medium

Read the debiased score next to the full score.

Reasoning

Read the debiased score next to the full score, as the authors recommend. A model whose score drops much more than others on the debiased subset is probably relying on answer patterns.

Confidence: medium

Known problems 6

  1. Many questions can be answered without the video

    A model that saw only the text of the questions learned their answer patterns. It gained 17.9 points without seeing any video.1192

    Details

    The original paper found blind models below chance overall, but object-size questions answerable above chance from common knowledge. The TsT study (COLM 2026, same lab) trained a text-only Qwen2-7B on VSI-Bench's own questions by cross-validation: its score rose from 24.7 to 42.7 (+17.9) without any video. Fine-tuning LLaVA-Video-7B on VSI-Train-10k, made with VSI-Bench's templates from the source datasets' training scans, raised its score with video from 36.7 to 57.1 and its score without video from 25.9 to 44.7. So much of what such training adds can come from answer patterns.

  2. Training data is built from the same scan datasets and templates

    Top models train on questions made with the same templates from the same scan datasets. Inferred71912+1

    Details

    The answer key is public, with no hidden test split. Several high-scoring models were trained on spatial question sets built from the training splits of the same scan datasets with VSI-Bench-style templates: VSI-590K (Cambrian-S, which follows VSI-Bench's recipe), VSI-Train-10k (TsT) and SpaceMind++'s 900K mix, which includes VSI-590K. The papers do not report test videos in these sets, but TsT shows that same-template training raises scores even without video. No study has checked scene overlap between the test videos and these training sets.

  3. Papers report different scores for the same model

    Different papers report GPT-5's score as 37.5, 52.9 and 55.0.5310+9

    Details

    GPT-5: 37.5 (InternVL3.5, VLMEvalKit), 52.9 (Gemini Robotics 1.5) and 55.0 (Spatial-TTT, Ouroboros-Spatial, SenseNova-SI model cards). Gemini 2.5 Pro: 43.4 (Vlaser), 51.1 (Gemini Robotics 1.5), 51.5 (Cambrian-S) and 53.5 (Spatial-TTT). Gemini 3 Pro: 56.0 (Spatial-TTT) and 57.9 (HY-Embodied-0.5 API run). Cambrian-S-7B: 67.5 in its paper (128 frames), 62.9 in Ouroboros-Spatial and 62.92 on the EASI board. SpaceMind: 69.6 in its own paper and 70.2 in SpaceMind++. Papers also differ in averaging (by task or by question), as the dataset card warns.

  4. The videos come with non-commercial terms from their sources

    The dataset is labelled Apache-2.0, but the source scans allow only non-commercial use. Inferred22122+1

    Details

    The Hugging Face card labels the dataset Apache-2.0 and serves the video files without a gate. The videos derive from ScanNet and ScanNet++, whose terms allow only non-commercial research and educational use after signing an agreement, and from ARKitScenes under Apple's licence. The card does not discuss these terms.

  5. Scores depend on how many frames a model sees

    With fewer video frames, some questions cannot be answered correctly. One model's score changed by 8.9 points with the number of frames.24813+2

    Details

    Models see a sample of video frames, and the number varies (8 to 128 across papers; Gemini models take the whole video at 1 frame per second). ReVSI checked how many questions keep a correct answer when frames are sampled evenly: with 16 frames, 28% of appearance-order and 54% of relative-distance questions; with 32 frames, 48% and 80%; with 64 frames, 70% and 92%. Scores move with frame count: Cambrian-S-7B scores 58.6 with 16 frames and 67.5 with 128 (SenseNova-SI paper), and Ouroboros-Spatial re-ran it at 32 frames and got 62.9. ReVSI recommends at least 64 frames.

  6. Some official answers are wrong

    The answers come from the labels of the source 3D scans. An audit (ReVSI) found errors in counting and size questions.2624

    Details

    VSI-Bench builds answers from the 3D annotations of the source scan datasets. A user listed counting questions whose official answers disagree with the video, for example 2 towels where 12 are visible (issue #40, 2025-09); a maintainer replied that source annotations contain errors and that filtering cannot ensure all answers are correct. ReVSI (ICML 2026) manually checked all object-counting, room-size and object-size questions. It reports notable errors and ambiguity in counting, room-size errors from noisy reconstructions, and physically implausible object sizes, and says similar errors are likely in other tasks.

Details

About

What it is
A benchmark with a fixed test set127
More

The paper presents VSI-Bench as a benchmark of fixed question-answer pairs with its own metrics and evaluation code.

Built by
New York University, Yale University, Stanford University127
More

Authors: Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, Saining Xie.

New York University · Four of six authors, including the senior author Saining Xie.1

Yale University · Rilyn Han.1

Stanford University · Li Fei-Fei.1

Released
December 2024, at CVPR 2025282927
More

arXiv v1 2024-12-18; data released on Hugging Face 2024-12-19. Accepted at CVPR 2025 as an oral paper.

Version
The full set has 5,130 questions. A debiased subset keeps 2,362.22829
More

Three configurations on Hugging Face: full (5,130 questions, the default), debiased (2,362 kept) and pruned (2,768 removed). Paper v2 (2025-07-02) updated open-model results to 32 frames.

VSI-Bench (tiny) · 400 questions, 50 per task, used for the human baseline.1

VSI-Bench-Debiased v1 · 2,362 of 5,130 questions kept by hand-tuned filters (added 2025-11-11). A 2026-08-09 note on the card says it predates the automated pruning method and is a separate evaluation set.219

Row order change · Rows were reordered on 2025-11-11 (commit d7cb1a3); the card says to join predictions on the question id.229

Last update
The dataset card changed in October 2026. The data last changed in November 2025.2930
More

Dataset card updated 2026-10-08 (debiased-subset description, lmms-eval task names, COLM 2026 citation). Last data change 2025-11-11 (VSI-Bench-Debiased added; row order changed). Code last changed 2025-08-05 (scene meta information released).

Status
Active. The dataset card was updated in October 2026. Inferred2930
More

The dataset card and debiased subset were updated through 2026-10-08. The evaluation code has not changed since 2025-08-05.

Recent changes are documentation; the last data change was 2025-11-11.

Setup

Runs in
Recorded data1
More

Models answer questions about recorded videos. Nothing is controlled.

Robot
No body1
More

Videos are walk-through scans of rooms. No body or robot is controlled.

Setting
Whole home, Office or lab, Industrial131
More

The paper lists residential spaces, professional settings (offices, labs) and industrial spaces (factories). A maintainer said the factory scene comes from ScanNet++ (issue #30).

Tasks
5,130 questions in 8 task types232
More

5,130 questions in 8 task types: object count 565, absolute distance 834, object size 953, room size 288, relative distance 710, relative direction 968 (easy 217, medium 378, hard 373), route plan 194, appearance order 618.

Counts from the Hugging Face dataset statistics for the full configuration (2026-10-10). The paper says 'over 5,000'. Four task types take numerical answers (2,640 questions) and four are multiple choice (2,490). Only route-plan questions were written by people; the rest come from templates.

Scenes
288 room videos13233
More

288 room videos from validation splits: ScanNet 88, ARKitScenes 150, ScanNet++ 50

288 is stated in the paper. The per-source counts are ours, from scene-name formats in the dataset statistics (inferred). Questions per source: ScanNet 2,071, ARKitScenes 1,601, ScanNet++ 1,458. The video download holds 512 videos, of which 288 are used (maintainer, issue #41).

Training data
None. Training sets built with the same templates exist elsewhere.2197
More

No training split. Training sets built with the same pipeline from the source datasets' training splits exist elsewhere: VSI-Train-10k (TsT paper) and VSI-590K (Cambrian-S).

Size
Videos of about 2 to 3 minutes3435
More

Videos average about 2 to 3 minutes (maintainer, issue #28). The three video archives total about 5.73 GB.

5,728,450,973 bytes summed by us from the Hugging Face file listing (arkitscenes.zip, scannet.zip, scannetpp.zip).

Changes at test
None stated. It is a test set only. Inferred12
More

A zero-shot test with no training split. Questions come from scenes in the validation splits of the source scan datasets.

Scoring and access

Scored by
Accuracy1
More

Multiple-choice questions use accuracy. Numerical questions use Mean Relative Accuracy (MRA), a partial-credit score; the taxonomy has no value for it.

Score
Average of 8 task scores12
More

Each task is scored on its own: multiple-choice questions by exact-match accuracy, numerical answers by MRA, the share of ten tolerance levels (from 50% down to 5% relative error) that the answer falls within. The headline is the plain average over the eight tasks (lmms-eval's default). The TsT paper averages over questions instead; the dataset card warns the two can differ.

Trials
One pass. The number of video frames varies by paper.12536+1
More

Each question is asked once with greedy decoding (temperature 0). In paper v2 the frames differ by model: 16 for GPT-4o, 32 for open models, and the whole video for Gemini, which samples 1 frame per second. Later papers use 16 to 128 frames.

Paper v1 to v2 · InternVL2 models moved from 8 to 32 frames between versions; InternVL2-8B went from 34.6 to 37.5.361

Frame sensitivity · SenseNova-SI paper: Cambrian-S-7B scores 58.6, 63.6, 66.4 and 67.5 with 16, 32, 64 and 128 frames.8

Who runs it
Each team tests its own model Inferred120
More

Each paper runs its own evaluation. The EASI community board runs models under one protocol.

Error bars
Not reported Inferred13419
More

Papers report single scores. The maintainers ran closed models several times and say overall scores were stable but individual answers were not (issue #28). The TsT paper reports rerun ranges only for its own probe (42.6 to 42.8).

Leaderboard
No official leaderboard. A community board (EASI) lists scores. Inferred203738+2
More

No official leaderboard. The EASI board (EvolvingLMMs-Lab) lists VSI-Bench and VSI-Bench-Debiased scores for 45 models, last updated 2026-07-01.

The project page, GitHub README and dataset card have no leaderboard. Embodied Arena also includes VSI-Bench (arXiv 2509.15273), but its site could not be read on 2026-10-10.

Code licence
Apache-2.01415
More

LICENSE file: Apache License 2.0. GitHub reports Apache-2.0.

Data licence
Apache-2.0, as labelled on the dataset card2
More

The dataset card says Apache-2.0. The videos come from scan datasets with non-commercial terms (see asset licence).

Asset licence
The source scans are under non-commercial terms212223
More

The videos come from ScanNet, ScanNet++ and ARKitScenes validation scans, each under its own terms.

The Hugging Face repository serves the videos without a gate under the Apache-2.0 label. The card does not discuss the source terms. Not legal advice.

ScanNet Terms of Use · Use only for non-commercial research and educational purposes; access after signing an agreement.21

ScanNet++ Terms of Use · Non-commercial research and educational purposes only; commercial use strictly prohibited; data may be shared with colleagues only after they agree to the terms.22

ARKitScenes licence (Apple) · A non-commercial licence, plus commercial terms for licensees whose products had fewer than 700 million monthly active users before August 2024; larger ones must ask Apple.23

Access
Open. The data is on Hugging Face.240
More

Ungated Hugging Face dataset; no registration. The source scan datasets themselves require signed agreements.

Commercial use
Not allowed Inferred22122+1
More

The card says Apache-2.0, but the videos derive from ScanNet and ScanNet++, whose terms allow only non-commercial research and educational use. Not legal advice.

Published at
CVPR 2025 (oral)2741
More

CVPR 2025 (oral, per the repository README; DOI 10.1109/CVPR52734.2025.00994)

Sources 41

  1. 1Thinking in Space, full text v2 (Sections 3 and 4, Figure 6, Appendix C)Paper · Jul 2025 · checked 10 Oct 2026
  2. 2nyu-visionx/VSI-Bench dataset cardRepository · 8 Oct 2026 · checked 10 Oct 2026
  3. 3Gemini Robotics 1.5 report, full text v3 (Appendix C.1, Table 19)Paper · Oct 2025 · checked 10 Oct 2026
  4. 4RoboBrain 2.0 Technical Report (Table 3: VSI-Bench)Paper · Jul 2025 · checked 10 Oct 2026
  5. 5InternVL3.5 report (Table 2 and Table 11)Paper · Aug 2025 · checked 10 Oct 2026
  6. 6Qwen3-VL Technical Report (Section 5.8)Paper · Nov 2025 · checked 10 Oct 2026
  7. 7Cambrian-S: Towards Spatial Supersensing in Video (Tables 5 and 6, VSI-590K)Paper · Nov 2025 · checked 10 Oct 2026
  8. 8Scaling Spatial Intelligence with Multimodal Foundation Models (SenseNova-SI), v4 (abstract, Table 2)Paper · Nov 2025 · checked 10 Oct 2026
  9. 9Vlaser (Table 1; Section 3.2 on transfer to robot control)Paper · Oct 2025 · checked 10 Oct 2026
  10. 10Spatial-TTT (Table 1: VSI-Bench)Paper · Mar 2026 · checked 10 Oct 2026
  11. 11HY-Embodied-0.5 report (Tables 1 and 2)Paper · Apr 2026 · checked 10 Oct 2026
  12. 12SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs (Table 1, training data)Paper · May 2026 · checked 10 Oct 2026
  13. 13Ouroboros-Spatial (Table 1: VSI-Bench at 32 frames)Paper · Jun 2026 · checked 10 Oct 2026
  14. 14thinking-in-space LICENSE (Apache 2.0)Repository · 19 Dec 2024 · checked 10 Oct 2026
  15. 15GitHub API: vision-x-nyu/thinking-in-space (stars, forks, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
  16. 16SenseNova-SI-1.2-InternVL3-8B model card (VSI 69.6; GPT-5 55.0)Repository · Dec 2025 · checked 10 Oct 2026
  17. 17SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning (Table 1)Paper · Nov 2025 · checked 10 Oct 2026
  18. 18Seeing Across Views: MV-RoboBench (Sections 2.4 and 4)Paper · Oct 2025 · checked 10 Oct 2026
  19. 19Benchmark Designers Should 'Train on the Test Set' to Expose Exploitable Non-Visual Shortcuts (TsT), v2Paper · Nov 2025 · checked 10 Oct 2026
  20. 20EASI leaderboard data API (VSI-Bench scores, 45 models)Leaderboard · 1 Jul 2026 · checked 10 Oct 2026
  21. 21ScanNet Terms of Use (PDF)Official site · unknown · checked 10 Oct 2026
  22. 22ScanNet++ Terms of Use (PDF)Official site · unknown · checked 10 Oct 2026
  23. 23ARKitScenes LICENSE (README points to it for dataset use)Repository · 2024 · checked 10 Oct 2026
  24. 24ReVSI: Rebuilding Visual Spatial Intelligence Evaluation (Section 3, Figure 2, Appendix A Table 7)Paper · Apr 2026 · checked 10 Oct 2026
  25. 25thinking-in-space issue #5: Number of frames used in evaluationRepository · 8 Jan 2025 · checked 10 Oct 2026
  26. 26thinking-in-space issue #40: Questionable ground truth in QA and metadataRepository · 26 Sep 2025 · checked 10 Oct 2026
  27. 27vision-x-nyu/thinking-in-space READMERepository · 5 Aug 2025 · checked 10 Oct 2026
  28. 28Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (arXiv abstract page)Paper · Dec 2024 · checked 10 Oct 2026
  29. 29Hugging Face commit list for nyu-visionx/VSI-BenchRepository · 8 Oct 2026 · checked 10 Oct 2026
  30. 30thinking-in-space commit historyRepository · 5 Aug 2025 · checked 10 Oct 2026
  31. 31thinking-in-space issue #30: the factory sceneRepository · 21 Jul 2025 · checked 10 Oct 2026
  32. 32Hugging Face dataset viewer statistics, full configuration (question types, sources, scenes)Index · 10 Oct 2026 · checked 10 Oct 2026
  33. 33thinking-in-space issue #41: number of videosRepository · 14 Oct 2025 · checked 10 Oct 2026
  34. 34thinking-in-space issue #28: Evaluation FramesRepository · 27 Jun 2025 · checked 10 Oct 2026
  35. 35Hugging Face file listing for nyu-visionx/VSI-Bench (file sizes)Index · 10 Oct 2026 · checked 10 Oct 2026
  36. 36Thinking in Space, full text v1 (Figure 6, Table 5 frame counts)Paper · Dec 2024 · checked 10 Oct 2026
  37. 37EvolvingLMMs-Lab/EASI README (benchmarks and leaderboard link)Repository · 1 Jul 2026 · checked 10 Oct 2026
  38. 38VSI-Bench project pageOfficial site · Dec 2024 · checked 10 Oct 2026
  39. 39Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AIPaper · Sep 2025 · checked 10 Oct 2026
  40. 40Hugging Face Hub API record for nyu-visionx/VSI-Bench (downloads, likes, gating)Index · 10 Oct 2026 · checked 10 Oct 2026
  41. 41Semantic Scholar API record for arXiv:2412.14171Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: VSI-Bench paper (v1 and v2), project page, repository issues; Gemini Robotics 1.5 report; Vlaser (2510.11027); MV-RoboBench (2510.19400); ReVSI (2604.24300); TsT (2511.04655); A2Eval (2602.01640); web searches on 2026-10-10 for studies that relate VSI-Bench to robot or VLA success. None found.

leaderboard: Project page, GitHub README, Hugging Face card (none official); EASI GitHub README and board API (read); Embodied Arena paper (read) and site (script-only page, not readable).

top_score: The papers and model cards under used_by and top_score; EASI board API (2026-07-01 snapshot); web searches on 2026-10-10 for VSI-Bench averages of 70 and above. SpaceMind++ (73.2) is the highest found.

license_assets: Dataset card, repository LICENSE, ScanNet README and Terms of Use PDF, ScanNet++ Terms of Use PDF, ARKitScenes LICENSE and README.

human_baseline (number of evaluators): Paper Section 4.1 and Appendix C.2; not stated.

contamination (scene overlap): Cambrian-S, TsT and SpaceMind++ data sections; web search on 2026-10-10. No study checks overlap between VSI-Bench test scenes and training sets.

Change history

  1. Created as a full entry from primary sources; no basic entry existed. Leads came from the niches, surveys and merged sweep records. Prior sweep claims checked: 288 videos from ScanNet, ScanNet++ and ARKitScenes validation splits, 5,130 questions, human 79% (79.2), Apache-2.0 card, 814 citations, 746 stars, 9,432 Hub downloads, debiased subset of 2,362, row-order change on 2025-11-11 and the 2026-08-09 provenance note were all confirmed. The sweep's 'best model about 33 points below humans' refers to the 2024 paper; the best result found now is 73.2. The sweep listed RoboBrain 2.0 as a user; confirmed.