ERQA
Embodied Reasoning Question Answering (ERQA), from the Gemini Robotics report
ERQA is a set of 400 multiple-choice questions about images from robots and first-person video. It scores how well vision-language models (AI models that read images and text) answer them, and no robot moves.12
What a score here does not tell you Inferred
- Whether a robot driven by the model will succeed at its task.No study has linked ERQA scores to robot task success.
- Whether a small gap between two models is real.With only 400 questions, scores have about ±5 points of sampling noise.
- How people would score on the same questions.No human baseline (a score from people taking the test) has been published.
Comparisons with real robots
Details
The Gemini Robotics report says zero-shot robot control 'is strongly correlated with better embodied understanding', based on one pair of models (Gemini 2.0 Flash and Gemini Robotics-ER); it does not report ERQA for Gemini Robotics-ER or relate ERQA scores to robot results. Gemini Robotics 1.5 reports ERQA (54.8 vs 47.5) and real-robot agent results for Gemini Robotics-ER 1.5 and Gemini 2.5 Flash, but ERQA is one of 15 benchmarks in its ER score and no comparison is made. Vlaser (ICLR 2026) found that gains on standard embodied-reasoning benchmarks, ERQA among them, did not carry over to robot control in simulation (SimplerEnv); that is a simulation result, not a real-robot comparison. ERIQ (arXiv 2512.24125) reports a positive correlation for its own benchmark, not for ERQA.
Our assessment Opinion
ERQA measures answers to questions about images. No study has linked it to robot success.
Reasoning
An ERQA score shows how often a vision-language model picks the right answer to questions about robot and first-person images. It is not evidence that a robot driven by that model will succeed. No study has linked ERQA scores to robot task success, and Vlaser found that gains on such benchmarks did not carry over to robot control in simulation.
Confidence: medium
Compare scores only within one report. Ignore gaps of less than about 5 points.
Reasoning
Compare ERQA numbers only within one report. Across reports, graders and prompts differ and the same model can move by up to 7.8 points. Within a report, gaps under about 5 points are inside the sampling noise of 400 questions.
Confidence: high
Most top scores come from Google's own runs on its own benchmark.
Reasoning
Most of the highest published scores come from Google, which built ERQA, chose the grader and runs competitors' models itself. An independent run gave Gemini 3 Pro 65.0 where Google's table gives 70.5. Read vendor charts as the vendor's own measurement.
Confidence: medium
Known problems 4
The same model gets different scores in different reports
The same model's score differs by up to 7.8 points between reports.385+8
Details
Reports use different graders, prompts and toolkits, and some copy numbers from other papers. GPT-5 scored 59.0 in Google's run (Gemini 2.5 Flash as grader) and 65.7 in InternVL3.5's run (VLMEvalKit). Gemini 3 Pro scored 70.5 in Google's table and 65.0 in HY-Embodied-0.5's API run (March 2026). Claude Opus 4.5 scored 51.3 in Google's table and 46.8 in the Qwen3.5 model card. Qwen3-VL-4B is reported at 47.3 (HY-Embodied-0.5), 41.2 (Ouroboros-Spatial), 40.0 (Xiaomi-Robotics-0) and 39.5 (LightNav-0). Google's July 2026 chart gives GPT 5.6 Sol 43.2, below the 59.0 Google reported for GPT-5 in 2025, and states no settings. Some papers also mislabel cited numbers: EO-1 lists 46.3 for Gemini 1.5 Flash, which the Gemini Robotics report gives to Gemini 2.0 Flash (1.5 Flash: 42.3), and InternVL3.5 attributes 48.3 to Gemini-2.5-Pro, which the report gives to Gemini 2.0 Pro Experimental.
The test is small and many models scored close together
With 400 questions, scores have about ±5 points of sampling noise. One study found most models scored between 40% and 50%.24254
Details
ERQA has 400 questions, and four categories have fewer than 40 (pointing 34, multi-view 37, task reasoning 38, other 14). A 95% sampling interval on 400 questions is about ±4.9 points at 50% accuracy (our calculation), so small gaps between models are within noise. MV-RoboBench (ICLR 2026) ran 11 models on ERQA and found most between 40% and 50%, with Qwen2.5-VL-7B (43.11%) less than 3 points behind GPT-4o (46.00%). Its authors judged ERQA to have low discriminative power and used another benchmark for their analysis. Later frontier results spread more widely (43.2 to 78.5 in Google's July 2026 chart).
Grader and prompt choices move scores
The grader (the program or model that marks answers) and the prompt change scores. A prompt that asks the model to reason step by step added up to 10.3 points.2636+4
Details
The released script counts an answer as correct only if the entire reply equals the answer letter, so any explanation counts as wrong. Google's Gemini Robotics 1.5 report used Gemini 2.5 Flash to grade answers; Qwen3-VL, EO-1 and InternVL3.5 used their own prompts or toolkits. In the first report, a chain-of-thought instruction raised Gemini 2.0 Pro Experimental from 48.3 to 54.8 and Claude 3.5 Sonnet from 35.5 to 45.8. The chain-of-thought code was not released; a user who followed the paper got 46.50% where the report gives 50.3% (issue #3, open since 2025-07-06, no reply).
The answers are public and some images come from robot training datasets
All questions and answers are public, and there is no hidden test set. No study has measured whether models saw them during training. Inferred2191+1
Details
All 400 questions and answers have been public on GitHub since March 2025 under CC BY 4.0, with no hidden test split, so later models may have seen them in training. No study has measured this. ERQA's images partly come from Open X-Embodiment and other public datasets that are used to train robot models. BAAI built its separate ERQA+ benchmark from newly annotated robot videos and lists reduced contamination as a goal, contrasting it with benchmarks that reuse earlier data.
Details
About
- What it is
- A benchmark with one fixed test set12
More
The report calls ERQA a benchmark. The repository ships one fixed test file with answers and an evaluation script.
- Built by
- Google DeepMind22930
More
Google DeepMind (Gemini Robotics Team)
README: released as part of Google DeepMind's Gemini Robotics release. The report is authored by the Gemini Robotics Team. All repository commits are by Ted Xiao, a report author.
- Released
- March 2025, with Gemini Robotics203031+2
More
Repository created 2025-03-07; test file added 2025-03-08; announced with Gemini Robotics on 2025-03-12; report on arXiv 2025-03-25.
The report is an arXiv technical report, not a peer-reviewed paper.
- Version
- One release with 400 questions233
More
One unversioned release: data/erqa.tfrecord (91,402,921 bytes) holding the 400 questions, an example loader and an evaluation script.
File size from the GitHub git-tree API. We did not download the file; the count of 400 comes from the README and the report.
- Last update
- March 2025. Nothing has changed since then.3020
More
Last commit 2025-03-12 (adds links to the blog post and report). The test file has not changed since 2025-03-08.
GitHub API pushed_at: 2025-03-12T19:53:12Z. No tags or releases.
- Status
- No changes to the repository since March 2025. It is still in active use. Inferred30344+1
More
No repository change since 2025-03-12. The five open issues (2025-06 to 2026-07) have no maintainer reply. Use keeps growing: Google reported ERQA again for Gemini Robotics ER 2 in July 2026.
Last commit is 19 months before 2026-10-10. The only reply on any open issue (#2) is from another user.
Setup
- Runs in
- Recorded data12
More
Models answer questions about still images. Nothing runs in a simulator or on a robot.
- Robot
- No body1
More
Images show robot arms and people's hands. The model being tested controls nothing.
- Setting
- A mix of robot scenes and first-person scenes Inferred1
More
Robot lab and household scenes, and first-person frames of people at work
Images are the authors' own or come from Open X-Embodiment (OXE), UMI Data, MECCANO, HoloAssist and EGTEA Gaze+ (report Section 2.1). The report gives no breakdown by source.
- Tasks
- 400 questions in 7 categories plus Other1242
More
400 multiple-choice questions in seven named categories plus Other: spatial reasoning 84, action reasoning 72, trajectory reasoning 66, state estimation 55, task reasoning 38, multi-view reasoning 37, pointing 34, other 14.
The category counts are printed in the report's Figure 4. We read them from the figure's source file (erqa_categories.svg) and matched each number to its label by position. They sum to 400.
- Training data
- None. ERQA is a test set only.233
More
No training data. The repository holds only the 400-question test file.
Scoring and access
- Score
- Percent of 400 questions answered correctly2631
More
Accuracy over the 400 questions. The released script counts an answer as correct only when the model's whole reply, with periods and spaces removed and case ignored, equals the answer letter. In the Gemini Robotics 1.5 report, Gemini 2.5 Flash graded the answers instead. The first report also gives results with a chain-of-thought instruction added to each question.
Under the released script, a model that explains its answer is marked wrong unless a grader or an answer-extraction step is added.
- Trials
- One pass over the 400 questions263
More
Each question is asked once. The script calls the model API at temperature 0. Gemini Robotics 1.5 used default thinking budgets and no tools.
Script default: 1 question · The script's --num_examples default is 1, while the README says the full benchmark is 400 examples. A run with default settings scores one question.262
GR 1.5 query window · Gemini 2.5 and GPT-5 models were queried between 2025-09-01 and 2025-09-20.3
- Who runs it
- Each team tests its own model Inferred134+1
More
Each model team runs ERQA itself. In its comparisons Google also runs other companies' models through their public APIs (for example GPT-5, Claude Opus 4.5, Opus 5 and GPT 5.6 Sol).
- Error bars
- Not reported Inferred134+1
More
All reports we opened give single accuracies without intervals. Gemini Robotics 1.5 averages 3 runs only in its thinking-budget plots (Fig. 16). With 400 questions, a 95% binomial interval is about ±4.9 points at 50% accuracy and ±4.2 points at 75% (our calculation: 1.96 x sqrt(p(1-p)/400)).
- Leaderboard
- Results appear only in papers. Inferred2136
More
No leaderboard in the repository or in Google's reports; results live in model reports. Embodied Arena (arXiv 2509.15273) describes community leaderboards that include ERQA, but its site is a script-only page we could not read on 2026-10-10. The EASI community board has no ERQA column.
- Code licence
- CC BY 4.01920
More
One LICENSE file (Creative Commons Attribution 4.0) at the repository root covers the scripts as well as the data. GitHub reports CC-BY-4.0.
- Data licence
- CC BY 4.019
More
The same LICENSE covers data/erqa.tfrecord.
- Asset licence
- Third-party images with mixed terms Inferred13738+2
More
Some images come from five public datasets with their own terms.
The report does not say which images come from which source, or on what basis third-party images are re-shared under CC BY 4.0. Not legal advice.
HoloAssist: CDLA-Permissive-2.0 · The HoloAssist site says the data is released under the CDLAv2 licence.37
Open X-Embodiment: CC BY 4.0 · OXE README: materials other than software are under CC BY 4.0. Its component datasets were not checked.38
UMI Data: not stated · The UMI repository's code is MIT; its README states no data licence.40
MECCANO: none found
EGTEA Gaze+: not checked
- Access
- Open. The test file is on GitHub.220
More
The test file is in the public GitHub repository. No registration.
- Commercial use
- Unclear Inferred19137+2
More
The ERQA licence (CC BY 4.0) allows commercial use with attribution. Some images come from third-party datasets whose terms we could not confirm (MECCANO, EGTEA Gaze+, UMI Data). Not legal advice.
- Published at
- arXiv technical report29
More
Introduced in the Gemini Robotics technical report (arXiv 2503.20020). No peer-reviewed venue.
Sources 40
- 1Gemini Robotics report, full text (Section 2.1, Tables 1 and 2, Figure 4)Paper · Mar 2025 · checked 10 Oct 2026
- 2embodiedreasoning/ERQA READMERepository · 12 Mar 2025 · checked 10 Oct 2026
- 3Gemini Robotics 1.5 report, full text v3 (Appendix C.1, Table 19)Paper · Oct 2025 · checked 10 Oct 2026
- 4Gemini Robotics ER 2 release post (chart: ER metrics comparison)Official blog · 30 Jul 2026 · checked 10 Oct 2026
- 5Gemini 3 Pro: the frontier of vision AI (benchmark table image)Official blog · 5 Dec 2025 · checked 10 Oct 2026
- 6Qwen3-VL Technical Report (Section 5.8)Paper · Nov 2025 · checked 10 Oct 2026
- 7Qwen3.5-397B-A17B model card (Spatial Intelligence table)Repository · Feb 2026 · checked 10 Oct 2026
- 8InternVL3.5 report (Table 2 and Table 11)Paper · Aug 2025 · checked 10 Oct 2026
- 9GLM-4.5V and GLM-4.1V-Thinking report, v6 (Table 2)Paper · Jul 2025 · checked 10 Oct 2026
- 10Seed1.8 Model Card (Table 2)Paper · Mar 2026 · checked 10 Oct 2026
- 11HY-Embodied-0.5 report (Tables 1 and 2)Paper · Apr 2026 · checked 10 Oct 2026
- 12Xiaomi-Robotics-0 report (Table 3)Paper · Feb 2026 · checked 10 Oct 2026
- 13EO-1 report, v5 (Table 2)Paper · Aug 2025 · checked 10 Oct 2026
- 14Wall-OSS-0.5 Technical Report (Table 7)Paper · May 2026 · checked 10 Oct 2026
- 15LightNav-0 (Table II)Paper · Aug 2026 · checked 10 Oct 2026
- 16Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning (Table 1, Section 3.2)Paper · Oct 2025 · checked 10 Oct 2026
- 17Ouroboros-Spatial (Table 2: ERQA)Paper · Jun 2026 · checked 10 Oct 2026
- 18Gemini Robotics ER 2 model cardOfficial site · 30 Jul 2026 · checked 10 Oct 2026
- 19ERQA LICENSE (Creative Commons Attribution 4.0)Repository · 12 Mar 2025 · checked 10 Oct 2026
- 20GitHub API: embodiedreasoning/ERQA (stars, forks, created, pushed, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
- 21ERIQ: Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-trainingPaper · Dec 2025 · checked 10 Oct 2026
- 22Gemini Robotics-ER 1.6 post (no ERQA number)Official blog · 14 Apr 2026 · checked 10 Oct 2026
- 23A2Eval: Agentic and Automated Evaluation for Embodied BrainPaper · Feb 2026 · checked 10 Oct 2026
- 24ERQA question categories, Figure 4 source file with countsPaper · Mar 2025 · checked 10 Oct 2026
- 25Seeing Across Views: MV-RoboBench (Appendix D.1: evaluation on ERQA)Paper · Oct 2025 · checked 10 Oct 2026
- 26ERQA evaluation harness (eval_harness.py)Repository · 10 Mar 2025 · checked 10 Oct 2026
- 27ERQA issue #3: Evaluation with CoTRepository · 6 Jul 2025 · checked 10 Oct 2026
- 28ERQA+: An Enhanced Benchmark on Embodied Reasoning (project page)Official site · 2025 · checked 10 Oct 2026
- 29Gemini Robotics: Bringing AI into the Physical World (arXiv abstract page)Paper · Mar 2025 · checked 10 Oct 2026
- 30ERQA commit historyRepository · 12 Mar 2025 · checked 10 Oct 2026
- 31Gemini Robotics brings AI into the physical world (launch post)Official blog · 12 Mar 2025 · checked 10 Oct 2026
- 32Gemini Robotics 1.5 arXiv abstract page (version history)Paper · Oct 2025 · checked 10 Oct 2026
- 33GitHub API: ERQA git tree (file list and sizes)Index · 10 Oct 2026 · checked 10 Oct 2026
- 34ERQA issues (six issues, none answered by a maintainer)Repository · Jul 2026 · checked 10 Oct 2026
- 35Gemini Robotics 2 brings whole body intelligence to robots (announcement)Official blog · 30 Jul 2026 · checked 10 Oct 2026
- 36Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AIPaper · Sep 2025 · checked 10 Oct 2026
- 37HoloAssist project site (data licence)Official site · 2024 · checked 10 Oct 2026
- 38Open X-Embodiment README (licence section)Repository · 2023 · checked 10 Oct 2026
- 39MECCANO repository (no licence file)Repository · 2021 · checked 10 Oct 2026
- 40Universal Manipulation Interface repository (code licence, data download notes)Repository · 2024 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: Gemini Robotics report (2503.20020v1, full text), Gemini Robotics 1.5 report (2510.03342v3, Sections 3 and 4, Appendix C), Gemini Robotics ER 1.6 post, ER 2 post and model card, Gemini Robotics 2 post; Vlaser (2510.11027), ERIQ (2512.24125), MV-RoboBench (2510.19400), A2Eval (2602.01640); web searches on 2026-10-10 for studies that relate ERQA or embodied-reasoning QA scores to robot or VLA success. None pairs ERQA scores with real-robot results.
human_baseline: Gemini Robotics report Section 2.1 and Tables 1 and 2; Gemini Robotics 1.5 Appendix C; ERQA README; ER 2 release post and model card. None reports a human score.
top_score: Google posts (Gemini Robotics, Gemini Robotics 1.5, Gemini 3 Pro vision, ER 1.6, ER 2), the model reports under used_by, the Qwen3.5 model card, and web searches for newer or higher ERQA scores. The ER 1.6 post has no ERQA number; its 72.5 comes from the ER 2 chart. The Gemini 3 Pro model card and evaluation PDF have no ERQA row; the 70.5 comes from Google's Gemini 3 Pro vision post.
license_assets: ERQA README and LICENSE; report Section 2.1; HoloAssist site; Open X-Embodiment README; UMI repository README; MECCANO GitHub repository and project site (no licence text found); EGTEA Gaze+ site (did not load).
leaderboard: ERQA README and reports; Embodied Arena paper (2509.15273) and site (script-only page); EASI community board API (no ERQA column).
citations: Semantic Scholar API by arXiv ID (record with 0 citations) and by title search (record with 482 citations, DOI 10.48550/arXiv.2503.20020).
region: Report author list, README, deepmind.google careers and about pages (no office list in text).
Change history
- Basic entry created (phase 1 re-verification).
- Full entry written from primary sources, starting from the basic entry and the frontier-labs inventory record; every fact re-checked. Added question-category counts (Figure 4 source file), scoring details from the released script, results from 15 model reports, the Google ER 2 and Gemini 3 Pro charts, issues on reporting, grading, test size and contamination, and licence terms of image sources. Corrections to earlier records: sim_to_real is now 'none-found' (was unknown); citations are 482 on the report's second Semantic Scholar record (the first shows 0); scale now lists categories. Confirmed prior leads: Gemini 2.5 Flash grading in GR 1.5, Feb 2025 results, ER 2 at 78.5%.
- Published as a full entry.