VSI-Bench
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
VSI-Bench is a set of 5,130 questions about videos of real rooms, used to test how well multimodal AI models (models that take in video and text) understand space. Models only answer questions, and no robot moves.12
What a score here does not tell you Inferred
- Whether a robot using the model will succeed at tasks.No study links VSI-Bench scores to robot results.
- Whether the model actually used the video.A model trained on the questions alone, without video, gained 17.9 points.
- How a score compares with scores in other papers.Papers use different numbers of video frames and different ways of averaging.
Comparisons with real robots
Details
MV-RoboBench (ICLR 2026) reports that strong scores on general single-view spatial benchmarks do not reliably carry over to robotic spatial questions in its own set; that compares question sets, not robot runs. Vlaser (ICLR 2026) found that gains on embodied-reasoning benchmarks, VSI-Bench among them, did not carry over to closed-loop robot control in simulation.
Our assessment Opinion
VSI-Bench measures how well a model answers spatial questions from video. No study links its scores to robot success.
Reasoning
VSI-Bench measures how well a model answers questions about the layout of real rooms from video. No study links its scores to robot task success, and in one study (Vlaser) gains on such benchmarks did not carry over to robot control in simulation.
Confidence: medium
Check the frame count, the averaging and the training data before comparing scores.
Reasoning
Before comparing two VSI-Bench numbers, check the frame count, the averaging method and whether the model was trained on same-template data. Each of these has moved scores by between 5 and 18 points.
Confidence: high
Humans score higher than models on layout questions, but several models beat them on size and distance.
Reasoning
The human score of 79.2 is an uneven reference. Humans scored 94 to 100 on layout and order tasks but 45.9 to 60.4 on size and distance estimates, where several 2026 models already score higher. It also rests on 50 questions per task.
Confidence: medium
Read the debiased score next to the full score.
Reasoning
Read the debiased score next to the full score, as the authors recommend. A model whose score drops much more than others on the debiased subset is probably relying on answer patterns.
Confidence: medium
Known problems 6
Many questions can be answered without the video
A model that saw only the text of the questions learned their answer patterns. It gained 17.9 points without seeing any video.1192
Details
The original paper found blind models below chance overall, but object-size questions answerable above chance from common knowledge. The TsT study (COLM 2026, same lab) trained a text-only Qwen2-7B on VSI-Bench's own questions by cross-validation: its score rose from 24.7 to 42.7 (+17.9) without any video. Fine-tuning LLaVA-Video-7B on VSI-Train-10k, made with VSI-Bench's templates from the source datasets' training scans, raised its score with video from 36.7 to 57.1 and its score without video from 25.9 to 44.7. So much of what such training adds can come from answer patterns.
Training data is built from the same scan datasets and templates
Top models train on questions made with the same templates from the same scan datasets. Inferred71912+1
Details
The answer key is public, with no hidden test split. Several high-scoring models were trained on spatial question sets built from the training splits of the same scan datasets with VSI-Bench-style templates: VSI-590K (Cambrian-S, which follows VSI-Bench's recipe), VSI-Train-10k (TsT) and SpaceMind++'s 900K mix, which includes VSI-590K. The papers do not report test videos in these sets, but TsT shows that same-template training raises scores even without video. No study has checked scene overlap between the test videos and these training sets.
Papers report different scores for the same model
Different papers report GPT-5's score as 37.5, 52.9 and 55.0.5310+9
Details
GPT-5: 37.5 (InternVL3.5, VLMEvalKit), 52.9 (Gemini Robotics 1.5) and 55.0 (Spatial-TTT, Ouroboros-Spatial, SenseNova-SI model cards). Gemini 2.5 Pro: 43.4 (Vlaser), 51.1 (Gemini Robotics 1.5), 51.5 (Cambrian-S) and 53.5 (Spatial-TTT). Gemini 3 Pro: 56.0 (Spatial-TTT) and 57.9 (HY-Embodied-0.5 API run). Cambrian-S-7B: 67.5 in its paper (128 frames), 62.9 in Ouroboros-Spatial and 62.92 on the EASI board. SpaceMind: 69.6 in its own paper and 70.2 in SpaceMind++. Papers also differ in averaging (by task or by question), as the dataset card warns.
The videos come with non-commercial terms from their sources
The dataset is labelled Apache-2.0, but the source scans allow only non-commercial use. Inferred22122+1
Details
The Hugging Face card labels the dataset Apache-2.0 and serves the video files without a gate. The videos derive from ScanNet and ScanNet++, whose terms allow only non-commercial research and educational use after signing an agreement, and from ARKitScenes under Apple's licence. The card does not discuss these terms.
Scores depend on how many frames a model sees
With fewer video frames, some questions cannot be answered correctly. One model's score changed by 8.9 points with the number of frames.24813+2
Details
Models see a sample of video frames, and the number varies (8 to 128 across papers; Gemini models take the whole video at 1 frame per second). ReVSI checked how many questions keep a correct answer when frames are sampled evenly: with 16 frames, 28% of appearance-order and 54% of relative-distance questions; with 32 frames, 48% and 80%; with 64 frames, 70% and 92%. Scores move with frame count: Cambrian-S-7B scores 58.6 with 16 frames and 67.5 with 128 (SenseNova-SI paper), and Ouroboros-Spatial re-ran it at 32 frames and got 62.9. ReVSI recommends at least 64 frames.
Some official answers are wrong
The answers come from the labels of the source 3D scans. An audit (ReVSI) found errors in counting and size questions.2624
Details
VSI-Bench builds answers from the 3D annotations of the source scan datasets. A user listed counting questions whose official answers disagree with the video, for example 2 towels where 12 are visible (issue #40, 2025-09); a maintainer replied that source annotations contain errors and that filtering cannot ensure all answers are correct. ReVSI (ICML 2026) manually checked all object-counting, room-size and object-size questions. It reports notable errors and ambiguity in counting, room-size errors from noisy reconstructions, and physically implausible object sizes, and says similar errors are likely in other tasks.
Details
About
- What it is
- A benchmark with a fixed test set127
More
The paper presents VSI-Bench as a benchmark of fixed question-answer pairs with its own metrics and evaluation code.
- Released
- December 2024, at CVPR 2025282927
More
arXiv v1 2024-12-18; data released on Hugging Face 2024-12-19. Accepted at CVPR 2025 as an oral paper.
- Version
- The full set has 5,130 questions. A debiased subset keeps 2,362.22829
More
Three configurations on Hugging Face: full (5,130 questions, the default), debiased (2,362 kept) and pruned (2,768 removed). Paper v2 (2025-07-02) updated open-model results to 32 frames.
VSI-Bench (tiny) · 400 questions, 50 per task, used for the human baseline.1
VSI-Bench-Debiased v1 · 2,362 of 5,130 questions kept by hand-tuned filters (added 2025-11-11). A 2026-08-09 note on the card says it predates the automated pruning method and is a separate evaluation set.219
Row order change · Rows were reordered on 2025-11-11 (commit d7cb1a3); the card says to join predictions on the question id.229
- Last update
- The dataset card changed in October 2026. The data last changed in November 2025.2930
More
Dataset card updated 2026-10-08 (debiased-subset description, lmms-eval task names, COLM 2026 citation). Last data change 2025-11-11 (VSI-Bench-Debiased added; row order changed). Code last changed 2025-08-05 (scene meta information released).
Setup
- Runs in
- Recorded data1
More
Models answer questions about recorded videos. Nothing is controlled.
- Robot
- No body1
More
Videos are walk-through scans of rooms. No body or robot is controlled.
- Setting
- Whole home, Office or lab, Industrial131
More
The paper lists residential spaces, professional settings (offices, labs) and industrial spaces (factories). A maintainer said the factory scene comes from ScanNet++ (issue #30).
- Tasks
- 5,130 questions in 8 task types232
More
5,130 questions in 8 task types: object count 565, absolute distance 834, object size 953, room size 288, relative distance 710, relative direction 968 (easy 217, medium 378, hard 373), route plan 194, appearance order 618.
Counts from the Hugging Face dataset statistics for the full configuration (2026-10-10). The paper says 'over 5,000'. Four task types take numerical answers (2,640 questions) and four are multiple choice (2,490). Only route-plan questions were written by people; the rest come from templates.
- Scenes
- 288 room videos13233
More
288 room videos from validation splits: ScanNet 88, ARKitScenes 150, ScanNet++ 50
288 is stated in the paper. The per-source counts are ours, from scene-name formats in the dataset statistics (inferred). Questions per source: ScanNet 2,071, ARKitScenes 1,601, ScanNet++ 1,458. The video download holds 512 videos, of which 288 are used (maintainer, issue #41).
- Training data
- None. Training sets built with the same templates exist elsewhere.2197
More
No training split. Training sets built with the same pipeline from the source datasets' training splits exist elsewhere: VSI-Train-10k (TsT paper) and VSI-590K (Cambrian-S).
Scoring and access
- Scored by
- Accuracy1
More
Multiple-choice questions use accuracy. Numerical questions use Mean Relative Accuracy (MRA), a partial-credit score; the taxonomy has no value for it.
- Score
- Average of 8 task scores12
More
Each task is scored on its own: multiple-choice questions by exact-match accuracy, numerical answers by MRA, the share of ten tolerance levels (from 50% down to 5% relative error) that the answer falls within. The headline is the plain average over the eight tasks (lmms-eval's default). The TsT paper averages over questions instead; the dataset card warns the two can differ.
- Trials
- One pass. The number of video frames varies by paper.12536+1
More
Each question is asked once with greedy decoding (temperature 0). In paper v2 the frames differ by model: 16 for GPT-4o, 32 for open models, and the whole video for Gemini, which samples 1 frame per second. Later papers use 16 to 128 frames.
Paper v1 to v2 · InternVL2 models moved from 8 to 32 frames between versions; InternVL2-8B went from 34.6 to 37.5.361
Frame sensitivity · SenseNova-SI paper: Cambrian-S-7B scores 58.6, 63.6, 66.4 and 67.5 with 16, 32, 64 and 128 frames.8
- Who runs it
- Each team tests its own model Inferred120
More
Each paper runs its own evaluation. The EASI community board runs models under one protocol.
- Error bars
- Not reported Inferred13419
More
Papers report single scores. The maintainers ran closed models several times and say overall scores were stable but individual answers were not (issue #28). The TsT paper reports rerun ranges only for its own probe (42.6 to 42.8).
- Leaderboard
- No official leaderboard. A community board (EASI) lists scores. Inferred203738+2
More
No official leaderboard. The EASI board (EvolvingLMMs-Lab) lists VSI-Bench and VSI-Bench-Debiased scores for 45 models, last updated 2026-07-01.
The project page, GitHub README and dataset card have no leaderboard. Embodied Arena also includes VSI-Bench (arXiv 2509.15273), but its site could not be read on 2026-10-10.
- Data licence
- Apache-2.0, as labelled on the dataset card2
More
The dataset card says Apache-2.0. The videos come from scan datasets with non-commercial terms (see asset licence).
- Asset licence
- The source scans are under non-commercial terms212223
More
The videos come from ScanNet, ScanNet++ and ARKitScenes validation scans, each under its own terms.
The Hugging Face repository serves the videos without a gate under the Apache-2.0 label. The card does not discuss the source terms. Not legal advice.
ScanNet Terms of Use · Use only for non-commercial research and educational purposes; access after signing an agreement.21
ScanNet++ Terms of Use · Non-commercial research and educational purposes only; commercial use strictly prohibited; data may be shared with colleagues only after they agree to the terms.22
ARKitScenes licence (Apple) · A non-commercial licence, plus commercial terms for licensees whose products had fewer than 700 million monthly active users before August 2024; larger ones must ask Apple.23
- Access
- Open. The data is on Hugging Face.240
More
Ungated Hugging Face dataset; no registration. The source scan datasets themselves require signed agreements.
Sources 41
- 1Thinking in Space, full text v2 (Sections 3 and 4, Figure 6, Appendix C)Paper · Jul 2025 · checked 10 Oct 2026
- 2nyu-visionx/VSI-Bench dataset cardRepository · 8 Oct 2026 · checked 10 Oct 2026
- 3Gemini Robotics 1.5 report, full text v3 (Appendix C.1, Table 19)Paper · Oct 2025 · checked 10 Oct 2026
- 4RoboBrain 2.0 Technical Report (Table 3: VSI-Bench)Paper · Jul 2025 · checked 10 Oct 2026
- 5InternVL3.5 report (Table 2 and Table 11)Paper · Aug 2025 · checked 10 Oct 2026
- 6Qwen3-VL Technical Report (Section 5.8)Paper · Nov 2025 · checked 10 Oct 2026
- 7Cambrian-S: Towards Spatial Supersensing in Video (Tables 5 and 6, VSI-590K)Paper · Nov 2025 · checked 10 Oct 2026
- 8Scaling Spatial Intelligence with Multimodal Foundation Models (SenseNova-SI), v4 (abstract, Table 2)Paper · Nov 2025 · checked 10 Oct 2026
- 9Vlaser (Table 1; Section 3.2 on transfer to robot control)Paper · Oct 2025 · checked 10 Oct 2026
- 10Spatial-TTT (Table 1: VSI-Bench)Paper · Mar 2026 · checked 10 Oct 2026
- 11HY-Embodied-0.5 report (Tables 1 and 2)Paper · Apr 2026 · checked 10 Oct 2026
- 12SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs (Table 1, training data)Paper · May 2026 · checked 10 Oct 2026
- 13Ouroboros-Spatial (Table 1: VSI-Bench at 32 frames)Paper · Jun 2026 · checked 10 Oct 2026
- 14thinking-in-space LICENSE (Apache 2.0)Repository · 19 Dec 2024 · checked 10 Oct 2026
- 15GitHub API: vision-x-nyu/thinking-in-space (stars, forks, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
- 16SenseNova-SI-1.2-InternVL3-8B model card (VSI 69.6; GPT-5 55.0)Repository · Dec 2025 · checked 10 Oct 2026
- 17SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning (Table 1)Paper · Nov 2025 · checked 10 Oct 2026
- 18Seeing Across Views: MV-RoboBench (Sections 2.4 and 4)Paper · Oct 2025 · checked 10 Oct 2026
- 19Benchmark Designers Should 'Train on the Test Set' to Expose Exploitable Non-Visual Shortcuts (TsT), v2Paper · Nov 2025 · checked 10 Oct 2026
- 20EASI leaderboard data API (VSI-Bench scores, 45 models)Leaderboard · 1 Jul 2026 · checked 10 Oct 2026
- 21ScanNet Terms of Use (PDF)Official site · unknown · checked 10 Oct 2026
- 22ScanNet++ Terms of Use (PDF)Official site · unknown · checked 10 Oct 2026
- 23ARKitScenes LICENSE (README points to it for dataset use)Repository · 2024 · checked 10 Oct 2026
- 24ReVSI: Rebuilding Visual Spatial Intelligence Evaluation (Section 3, Figure 2, Appendix A Table 7)Paper · Apr 2026 · checked 10 Oct 2026
- 25thinking-in-space issue #5: Number of frames used in evaluationRepository · 8 Jan 2025 · checked 10 Oct 2026
- 26thinking-in-space issue #40: Questionable ground truth in QA and metadataRepository · 26 Sep 2025 · checked 10 Oct 2026
- 27vision-x-nyu/thinking-in-space READMERepository · 5 Aug 2025 · checked 10 Oct 2026
- 28Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (arXiv abstract page)Paper · Dec 2024 · checked 10 Oct 2026
- 29Hugging Face commit list for nyu-visionx/VSI-BenchRepository · 8 Oct 2026 · checked 10 Oct 2026
- 30thinking-in-space commit historyRepository · 5 Aug 2025 · checked 10 Oct 2026
- 31thinking-in-space issue #30: the factory sceneRepository · 21 Jul 2025 · checked 10 Oct 2026
- 32Hugging Face dataset viewer statistics, full configuration (question types, sources, scenes)Index · 10 Oct 2026 · checked 10 Oct 2026
- 33thinking-in-space issue #41: number of videosRepository · 14 Oct 2025 · checked 10 Oct 2026
- 34thinking-in-space issue #28: Evaluation FramesRepository · 27 Jun 2025 · checked 10 Oct 2026
- 35Hugging Face file listing for nyu-visionx/VSI-Bench (file sizes)Index · 10 Oct 2026 · checked 10 Oct 2026
- 36Thinking in Space, full text v1 (Figure 6, Table 5 frame counts)Paper · Dec 2024 · checked 10 Oct 2026
- 37EvolvingLMMs-Lab/EASI README (benchmarks and leaderboard link)Repository · 1 Jul 2026 · checked 10 Oct 2026
- 38VSI-Bench project pageOfficial site · Dec 2024 · checked 10 Oct 2026
- 39Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AIPaper · Sep 2025 · checked 10 Oct 2026
- 40Hugging Face Hub API record for nyu-visionx/VSI-Bench (downloads, likes, gating)Index · 10 Oct 2026 · checked 10 Oct 2026
- 41Semantic Scholar API record for arXiv:2412.14171Index · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: VSI-Bench paper (v1 and v2), project page, repository issues; Gemini Robotics 1.5 report; Vlaser (2510.11027); MV-RoboBench (2510.19400); ReVSI (2604.24300); TsT (2511.04655); A2Eval (2602.01640); web searches on 2026-10-10 for studies that relate VSI-Bench to robot or VLA success. None found.
leaderboard: Project page, GitHub README, Hugging Face card (none official); EASI GitHub README and board API (read); Embodied Arena paper (read) and site (script-only page, not readable).
top_score: The papers and model cards under used_by and top_score; EASI board API (2026-07-01 snapshot); web searches on 2026-10-10 for VSI-Bench averages of 70 and above. SpaceMind++ (73.2) is the highest found.
license_assets: Dataset card, repository LICENSE, ScanNet README and Terms of Use PDF, ScanNet++ Terms of Use PDF, ARKitScenes LICENSE and README.
human_baseline (number of evaluators): Paper Section 4.1 and Appendix C.2; not stated.
contamination (scene overlap): Cambrian-S, TsT and SpaceMind++ data sections; web search on 2026-10-10. No study checks overlap between VSI-Bench test scenes and training sets.
Change history
- Created as a full entry from primary sources; no basic entry existed. Leads came from the niches, surveys and merged sweep records. Prior sweep claims checked: 288 videos from ScanNet, ScanNet++ and ARKitScenes validation splits, 5,130 questions, human 79% (79.2), Apache-2.0 card, 814 citations, 746 stars, 9,432 Hub downloads, debiased subset of 2,362, row-order change on 2025-11-11 and the 2026-08-09 provenance note were all confirmed. The sweep's 'best model about 33 points below humans' refers to the 2024 paper; the best result found now is 73.2. The sweep listed RoboBrain 2.0 as a user; confirmed.