ERQA

Embodied Reasoning Question Answering (ERQA), from the Gemini Robotics report

How to read this picture

ERQA is a set of 400 multiple-choice questions about images from robots and first-person video. It scores how well vision-language models (AI models that read images and text) answer them, and no robot moves.12

Sources
Last checked 10 Oct 2026Full entry60 of 75 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data12
Checked against real robots
Not checked
Skill
Reasoning
Robot
No body1
Used by
15 model reports, from March 2025 to September 2026134+14
482 citations
Licence
CC BY 4.01920
Commercial use: unclear

What a score here does not tell you Inferred

  1. Whether a robot driven by the model will succeed at its task.No study has linked ERQA scores to robot task success.
  2. Whether a small gap between two models is real.With only 400 questions, scores have about ±5 points of sampling noise.
  3. How people would score on the same questions.No human baseline (a score from people taking the test) has been published.
ChartPublished scores over time
95% AND ABOVE40506070809010020252026Gemini 2.0 Pro Experimental · 48.3% · 2025-03InternVL3.5-241B-A28B · 46.8% · 2025-08GPT-5 (run by Google) · 59% · 2025-10Gemini Robotics-ER 1.5 · 54.8% · 2025-10Qwen3-VL-235B-A22B · 52.5% · 2025-11Gemini 3 Pro · 70.5% · 2025-12Qwen3.5-397B-A17B · 67.5% · 2026-02Seed1.8 · 58.8% · 2026-03HY-Embodied-0.5 MoE-A32B · 62.3% · 2026-04Gemini 3.6 Flash · 73% · 2026-07Gemini Robotics ER 2 · 78.5% · 2026-07Gemini 2.0 Pro Experimental 48.3%Gemini Robotics ER 2 78.5%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.183+6

Comparisons with real robots

Not checked Inferred1316+3

Details

The Gemini Robotics report says zero-shot robot control 'is strongly correlated with better embodied understanding', based on one pair of models (Gemini 2.0 Flash and Gemini Robotics-ER); it does not report ERQA for Gemini Robotics-ER or relate ERQA scores to robot results. Gemini Robotics 1.5 reports ERQA (54.8 vs 47.5) and real-robot agent results for Gemini Robotics-ER 1.5 and Gemini 2.5 Flash, but ERQA is one of 15 benchmarks in its ER score and no comparison is made. Vlaser (ICLR 2026) found that gains on standard embodied-reasoning benchmarks, ERQA among them, did not carry over to robot control in simulation (SimplerEnv); that is a simulation result, not a real-robot comparison. ERIQ (arXiv 2512.24125) reports a positive correlation for its own benchmark, not for ERQA.

Our assessment Opinion

ERQA measures answers to questions about images. No study has linked it to robot success.

Reasoning

An ERQA score shows how often a vision-language model picks the right answer to questions about robot and first-person images. It is not evidence that a robot driven by that model will succeed. No study has linked ERQA scores to robot task success, and Vlaser found that gains on such benchmarks did not carry over to robot control in simulation.

Confidence: medium

Compare scores only within one report. Ignore gaps of less than about 5 points.

Reasoning

Compare ERQA numbers only within one report. Across reports, graders and prompts differ and the same model can move by up to 7.8 points. Within a report, gaps under about 5 points are inside the sampling noise of 400 questions.

Confidence: high

Most top scores come from Google's own runs on its own benchmark.

Reasoning

Most of the highest published scores come from Google, which built ERQA, chose the grader and runs competitors' models itself. An independent run gave Gemini 3 Pro 65.0 where Google's table gives 70.5. Read vendor charts as the vendor's own measurement.

Confidence: medium

Known problems 4

  1. The same model gets different scores in different reports

    The same model's score differs by up to 7.8 points between reports.385+8

    Details

    Reports use different graders, prompts and toolkits, and some copy numbers from other papers. GPT-5 scored 59.0 in Google's run (Gemini 2.5 Flash as grader) and 65.7 in InternVL3.5's run (VLMEvalKit). Gemini 3 Pro scored 70.5 in Google's table and 65.0 in HY-Embodied-0.5's API run (March 2026). Claude Opus 4.5 scored 51.3 in Google's table and 46.8 in the Qwen3.5 model card. Qwen3-VL-4B is reported at 47.3 (HY-Embodied-0.5), 41.2 (Ouroboros-Spatial), 40.0 (Xiaomi-Robotics-0) and 39.5 (LightNav-0). Google's July 2026 chart gives GPT 5.6 Sol 43.2, below the 59.0 Google reported for GPT-5 in 2025, and states no settings. Some papers also mislabel cited numbers: EO-1 lists 46.3 for Gemini 1.5 Flash, which the Gemini Robotics report gives to Gemini 2.0 Flash (1.5 Flash: 42.3), and InternVL3.5 attributes 48.3 to Gemini-2.5-Pro, which the report gives to Gemini 2.0 Pro Experimental.

  2. The test is small and many models scored close together

    With 400 questions, scores have about ±5 points of sampling noise. One study found most models scored between 40% and 50%.24254

    Details

    ERQA has 400 questions, and four categories have fewer than 40 (pointing 34, multi-view 37, task reasoning 38, other 14). A 95% sampling interval on 400 questions is about ±4.9 points at 50% accuracy (our calculation), so small gaps between models are within noise. MV-RoboBench (ICLR 2026) ran 11 models on ERQA and found most between 40% and 50%, with Qwen2.5-VL-7B (43.11%) less than 3 points behind GPT-4o (46.00%). Its authors judged ERQA to have low discriminative power and used another benchmark for their analysis. Later frontier results spread more widely (43.2 to 78.5 in Google's July 2026 chart).

  3. Grader and prompt choices move scores

    The grader (the program or model that marks answers) and the prompt change scores. A prompt that asks the model to reason step by step added up to 10.3 points.2636+4

    Details

    The released script counts an answer as correct only if the entire reply equals the answer letter, so any explanation counts as wrong. Google's Gemini Robotics 1.5 report used Gemini 2.5 Flash to grade answers; Qwen3-VL, EO-1 and InternVL3.5 used their own prompts or toolkits. In the first report, a chain-of-thought instruction raised Gemini 2.0 Pro Experimental from 48.3 to 54.8 and Claude 3.5 Sonnet from 35.5 to 45.8. The chain-of-thought code was not released; a user who followed the paper got 46.50% where the report gives 50.3% (issue #3, open since 2025-07-06, no reply).

  4. The answers are public and some images come from robot training datasets

    All questions and answers are public, and there is no hidden test set. No study has measured whether models saw them during training. Inferred2191+1

    Details

    All 400 questions and answers have been public on GitHub since March 2025 under CC BY 4.0, with no hidden test split, so later models may have seen them in training. No study has measured this. ERQA's images partly come from Open X-Embodiment and other public datasets that are used to train robot models. BAAI built its separate ERQA+ benchmark from newly annotated robot videos and lists reduced contamination as a goal, contrasting it with benchmarks that reuse earlier data.

Details

About

What it is
A benchmark with one fixed test set12
More

The report calls ERQA a benchmark. The repository ships one fixed test file with answers and an evaluation script.

Built by
Google DeepMind22930
More

Google DeepMind (Gemini Robotics Team)

README: released as part of Google DeepMind's Gemini Robotics release. The report is authored by the Gemini Robotics Team. All repository commits are by Ted Xiao, a report author.

Released
March 2025, with Gemini Robotics203031+2
More

Repository created 2025-03-07; test file added 2025-03-08; announced with Gemini Robotics on 2025-03-12; report on arXiv 2025-03-25.

The report is an arXiv technical report, not a peer-reviewed paper.

Version
One release with 400 questions233
More

One unversioned release: data/erqa.tfrecord (91,402,921 bytes) holding the 400 questions, an example loader and an evaluation script.

File size from the GitHub git-tree API. We did not download the file; the count of 400 comes from the README and the report.

Last update
March 2025. Nothing has changed since then.3020
More

Last commit 2025-03-12 (adds links to the blog post and report). The test file has not changed since 2025-03-08.

GitHub API pushed_at: 2025-03-12T19:53:12Z. No tags or releases.

Status
No changes to the repository since March 2025. It is still in active use. Inferred30344+1
More

No repository change since 2025-03-12. The five open issues (2025-06 to 2026-07) have no maintainer reply. Use keeps growing: Google reported ERQA again for Gemini Robotics ER 2 in July 2026.

Last commit is 19 months before 2026-10-10. The only reply on any open issue (#2) is from another user.

Setup

Runs in
Recorded data12
More

Models answer questions about still images. Nothing runs in a simulator or on a robot.

Robot
No body1
More

Images show robot arms and people's hands. The model being tested controls nothing.

Setting
A mix of robot scenes and first-person scenes Inferred1
More

Robot lab and household scenes, and first-person frames of people at work

Images are the authors' own or come from Open X-Embodiment (OXE), UMI Data, MECCANO, HoloAssist and EGTEA Gaze+ (report Section 2.1). The report gives no breakdown by source.

Tasks
400 questions in 7 categories plus Other1242
More

400 multiple-choice questions in seven named categories plus Other: spatial reasoning 84, action reasoning 72, trajectory reasoning 66, state estimation 55, task reasoning 38, multi-view reasoning 37, pointing 34, other 14.

The category counts are printed in the report's Figure 4. We read them from the figure's source file (erqa_categories.svg) and matched each number to its label by position. They sum to 400.

Training data
None. ERQA is a test set only.233
More

No training data. The repository holds only the 400-question test file.

Size
28% of questions use several images12
More

Each question mixes text and one or more images. 28% of questions have more than one image, and the report says these tend to be harder. Answers are a single letter, A to D (README).

Changes at test
None stated. ERQA is a test set only. Inferred12
More

ERQA is one fixed test set with no training split. The report describes no controlled change between training and test conditions.

Scoring and access

Scored by
Accuracy126
More

Multiple-choice accuracy (report Table 1).

Score
Percent of 400 questions answered correctly2631
More

Accuracy over the 400 questions. The released script counts an answer as correct only when the model's whole reply, with periods and spaces removed and case ignored, equals the answer letter. In the Gemini Robotics 1.5 report, Gemini 2.5 Flash graded the answers instead. The first report also gives results with a chain-of-thought instruction added to each question.

Under the released script, a model that explains its answer is marked wrong unless a grader or an answer-extraction step is added.

Trials
One pass over the 400 questions263
More

Each question is asked once. The script calls the model API at temperature 0. Gemini Robotics 1.5 used default thinking budgets and no tools.

Script default: 1 question · The script's --num_examples default is 1, while the README says the full benchmark is 400 examples. A run with default settings scores one question.262

GR 1.5 query window · Gemini 2.5 and GPT-5 models were queried between 2025-09-01 and 2025-09-20.3

Who runs it
Each team tests its own model Inferred134+1
More

Each model team runs ERQA itself. In its comparisons Google also runs other companies' models through their public APIs (for example GPT-5, Claude Opus 4.5, Opus 5 and GPT 5.6 Sol).

Error bars
Not reported Inferred134+1
More

All reports we opened give single accuracies without intervals. Gemini Robotics 1.5 averages 3 runs only in its thinking-budget plots (Fig. 16). With 400 questions, a 95% binomial interval is about ±4.9 points at 50% accuracy and ±4.2 points at 75% (our calculation: 1.96 x sqrt(p(1-p)/400)).

Leaderboard
Results appear only in papers. Inferred2136
More

No leaderboard in the repository or in Google's reports; results live in model reports. Embodied Arena (arXiv 2509.15273) describes community leaderboards that include ERQA, but its site is a script-only page we could not read on 2026-10-10. The EASI community board has no ERQA column.

Code licence
CC BY 4.01920
More

One LICENSE file (Creative Commons Attribution 4.0) at the repository root covers the scripts as well as the data. GitHub reports CC-BY-4.0.

Data licence
CC BY 4.019
More

The same LICENSE covers data/erqa.tfrecord.

Asset licence
Third-party images with mixed terms Inferred13738+2
More

Some images come from five public datasets with their own terms.

The report does not say which images come from which source, or on what basis third-party images are re-shared under CC BY 4.0. Not legal advice.

HoloAssist: CDLA-Permissive-2.0 · The HoloAssist site says the data is released under the CDLAv2 licence.37

Open X-Embodiment: CC BY 4.0 · OXE README: materials other than software are under CC BY 4.0. Its component datasets were not checked.38

UMI Data: not stated · The UMI repository's code is MIT; its README states no data licence.40

MECCANO: none found

EGTEA Gaze+: not checked

Access
Open. The test file is on GitHub.220
More

The test file is in the public GitHub repository. No registration.

Commercial use
Unclear Inferred19137+2
More

The ERQA licence (CC BY 4.0) allows commercial use with attribution. Some images come from third-party datasets whose terms we could not confirm (MECCANO, EGTEA Gaze+, UMI Data). Not legal advice.

Published at
arXiv technical report29
More

Introduced in the Gemini Robotics technical report (arXiv 2503.20020). No peer-reviewed venue.

Sources 40

  1. 1Gemini Robotics report, full text (Section 2.1, Tables 1 and 2, Figure 4)Paper · Mar 2025 · checked 10 Oct 2026
  2. 2embodiedreasoning/ERQA READMERepository · 12 Mar 2025 · checked 10 Oct 2026
  3. 3Gemini Robotics 1.5 report, full text v3 (Appendix C.1, Table 19)Paper · Oct 2025 · checked 10 Oct 2026
  4. 4Gemini Robotics ER 2 release post (chart: ER metrics comparison)Official blog · 30 Jul 2026 · checked 10 Oct 2026
  5. 5Gemini 3 Pro: the frontier of vision AI (benchmark table image)Official blog · 5 Dec 2025 · checked 10 Oct 2026
  6. 6Qwen3-VL Technical Report (Section 5.8)Paper · Nov 2025 · checked 10 Oct 2026
  7. 7Qwen3.5-397B-A17B model card (Spatial Intelligence table)Repository · Feb 2026 · checked 10 Oct 2026
  8. 8InternVL3.5 report (Table 2 and Table 11)Paper · Aug 2025 · checked 10 Oct 2026
  9. 9GLM-4.5V and GLM-4.1V-Thinking report, v6 (Table 2)Paper · Jul 2025 · checked 10 Oct 2026
  10. 10Seed1.8 Model Card (Table 2)Paper · Mar 2026 · checked 10 Oct 2026
  11. 11HY-Embodied-0.5 report (Tables 1 and 2)Paper · Apr 2026 · checked 10 Oct 2026
  12. 12Xiaomi-Robotics-0 report (Table 3)Paper · Feb 2026 · checked 10 Oct 2026
  13. 13EO-1 report, v5 (Table 2)Paper · Aug 2025 · checked 10 Oct 2026
  14. 14Wall-OSS-0.5 Technical Report (Table 7)Paper · May 2026 · checked 10 Oct 2026
  15. 15LightNav-0 (Table II)Paper · Aug 2026 · checked 10 Oct 2026
  16. 16Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning (Table 1, Section 3.2)Paper · Oct 2025 · checked 10 Oct 2026
  17. 17Ouroboros-Spatial (Table 2: ERQA)Paper · Jun 2026 · checked 10 Oct 2026
  18. 18Gemini Robotics ER 2 model cardOfficial site · 30 Jul 2026 · checked 10 Oct 2026
  19. 19ERQA LICENSE (Creative Commons Attribution 4.0)Repository · 12 Mar 2025 · checked 10 Oct 2026
  20. 20GitHub API: embodiedreasoning/ERQA (stars, forks, created, pushed, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
  21. 21ERIQ: Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-trainingPaper · Dec 2025 · checked 10 Oct 2026
  22. 22Gemini Robotics-ER 1.6 post (no ERQA number)Official blog · 14 Apr 2026 · checked 10 Oct 2026
  23. 23A2Eval: Agentic and Automated Evaluation for Embodied BrainPaper · Feb 2026 · checked 10 Oct 2026
  24. 24ERQA question categories, Figure 4 source file with countsPaper · Mar 2025 · checked 10 Oct 2026
  25. 25Seeing Across Views: MV-RoboBench (Appendix D.1: evaluation on ERQA)Paper · Oct 2025 · checked 10 Oct 2026
  26. 26ERQA evaluation harness (eval_harness.py)Repository · 10 Mar 2025 · checked 10 Oct 2026
  27. 27ERQA issue #3: Evaluation with CoTRepository · 6 Jul 2025 · checked 10 Oct 2026
  28. 28ERQA+: An Enhanced Benchmark on Embodied Reasoning (project page)Official site · 2025 · checked 10 Oct 2026
  29. 29Gemini Robotics: Bringing AI into the Physical World (arXiv abstract page)Paper · Mar 2025 · checked 10 Oct 2026
  30. 30ERQA commit historyRepository · 12 Mar 2025 · checked 10 Oct 2026
  31. 31Gemini Robotics brings AI into the physical world (launch post)Official blog · 12 Mar 2025 · checked 10 Oct 2026
  32. 32Gemini Robotics 1.5 arXiv abstract page (version history)Paper · Oct 2025 · checked 10 Oct 2026
  33. 33GitHub API: ERQA git tree (file list and sizes)Index · 10 Oct 2026 · checked 10 Oct 2026
  34. 34ERQA issues (six issues, none answered by a maintainer)Repository · Jul 2026 · checked 10 Oct 2026
  35. 35Gemini Robotics 2 brings whole body intelligence to robots (announcement)Official blog · 30 Jul 2026 · checked 10 Oct 2026
  36. 36Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AIPaper · Sep 2025 · checked 10 Oct 2026
  37. 37HoloAssist project site (data licence)Official site · 2024 · checked 10 Oct 2026
  38. 38Open X-Embodiment README (licence section)Repository · 2023 · checked 10 Oct 2026
  39. 39MECCANO repository (no licence file)Repository · 2021 · checked 10 Oct 2026
  40. 40Universal Manipulation Interface repository (code licence, data download notes)Repository · 2024 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: Gemini Robotics report (2503.20020v1, full text), Gemini Robotics 1.5 report (2510.03342v3, Sections 3 and 4, Appendix C), Gemini Robotics ER 1.6 post, ER 2 post and model card, Gemini Robotics 2 post; Vlaser (2510.11027), ERIQ (2512.24125), MV-RoboBench (2510.19400), A2Eval (2602.01640); web searches on 2026-10-10 for studies that relate ERQA or embodied-reasoning QA scores to robot or VLA success. None pairs ERQA scores with real-robot results.

human_baseline: Gemini Robotics report Section 2.1 and Tables 1 and 2; Gemini Robotics 1.5 Appendix C; ERQA README; ER 2 release post and model card. None reports a human score.

top_score: Google posts (Gemini Robotics, Gemini Robotics 1.5, Gemini 3 Pro vision, ER 1.6, ER 2), the model reports under used_by, the Qwen3.5 model card, and web searches for newer or higher ERQA scores. The ER 1.6 post has no ERQA number; its 72.5 comes from the ER 2 chart. The Gemini 3 Pro model card and evaluation PDF have no ERQA row; the 70.5 comes from Google's Gemini 3 Pro vision post.

license_assets: ERQA README and LICENSE; report Section 2.1; HoloAssist site; Open X-Embodiment README; UMI repository README; MECCANO GitHub repository and project site (no licence text found); EGTEA Gaze+ site (did not load).

leaderboard: ERQA README and reports; Embodied Arena paper (2509.15273) and site (script-only page); EASI community board API (no ERQA column).

citations: Semantic Scholar API by arXiv ID (record with 0 citations) and by title search (record with 482 citations, DOI 10.48550/arXiv.2503.20020).

region: Report author list, README, deepmind.google careers and about pages (no office list in text).

Change history

  1. Basic entry created (phase 1 re-verification).
  2. Full entry written from primary sources, starting from the basic entry and the frontier-labs inventory record; every fact re-checked. Added question-category counts (Figure 4 source file), scoring details from the released script, results from 15 model reports, the Google ER 2 and Gemini 3 Pro charts, issues on reporting, grading, test size and contamination, and licence terms of image sources. Corrections to earlier records: sim_to_real is now 'none-found' (was unknown); citations are 482 on the report's second Semantic Scholar record (the first shows 0); scale now lists categories. Confirmed prior leads: Gemini 2.5 Flash grading in GR 1.5, Feb 2025 results, ER 2 at 78.5%.
  3. Published as a full entry.