ASIMOV
Generating Robot Constitutions & Benchmarks for Semantic Safety (ASIMOV v1); Can AI Perceive Physical Danger and Intervene? (ASIMOV-2.0)
ASIMOV is a set of question sets from Google DeepMind. It tests whether AI models recognise physical danger the way human raters do.123
What a score here does not tell you Inferred
- Whether a robot run by the model acts safely.Models answer questions, and no robot moves.
- How well a model judges danger in real camera footage.The images and videos in the test are AI-generated.
- How a model does on questions it has not seen before.The answers are public, and there is no hidden test set.
Comparisons with real robots
Details
The scores are already agreement with human judges, so the open question is whether answering well predicts acting safely. No paired study found (see searched).
Our assessment Opinion
ASIMOV tests how a model judges danger. It does not test whether a robot acts safely.
Reasoning
ASIMOV measures whether a model's answers about physical danger match human raters. A high score shows the model recognises risks described in text or shown in generated media. It does not show that a robot run by the model will act safely.
Confidence: high
Text questions are nearly solved. Video questions and questions about robot limits are not.
Reasoning
The most useful information is the gap between question types. Frontier models (the most capable current models) score close to the maximum on text risk questions. They score far below it on video timing questions and on questions about robot limits.
Confidence: medium
Most published results come from Google. Compare scores across companies with care.
Reasoning
Most published results come from Google, which also trains its robot models on similar data, and the answers are public. Compare scores across companies with care.
Confidence: medium
Known problems 7
Recognising a danger is different from acting safely
A model that recognises danger in a question may still fail to act safely.1311
Details
ASIMOV asks models to judge scenarios; it does not test whether they act safely. SafetyALFRED (ACL 2026 Findings) found models that recognised kitchen hazards in question form often failed to deal with them when planning, and concluded that static question-answer evaluations are insufficient for physical safety; it names ASIMOV as such a benchmark but did not use its data. Another 2026 paper describes ASIMOV as testing whether models judge scenarios as safe 'without executing them'.
Images and videos are AI-generated
The scenes are generated, so results on real camera footage have not been tested.129+2
Details
v1 images are real frames edited by Imagen 3, and the 2.0 videos and constraint images are generated by Veo 3 and Imagen 3. The 2.0 paper filtered videos that were not photorealistic or broke physics. VLESA notes that robustness on real egocentric video 'remains unverified'. HomeSafe-Bench says ASIMOV-v2 targets general hazards and lacks diversity for household agent behaviours, and a 2026 collision-grounding paper notes it has no human-robot co-presence, depth or physics labels.
Answers are public and the builder trains on similar data
The answers are public. Google trains its own models on data made the same way. Inferred16176+2
Details
All released evaluation sets include their answers, and there is no hidden test set, so later models may have seen them in training. No study has measured this. Google's Gemini Robotics report says Gemini Robotics-ER models are post-trained on ASIMOV-type instances. The best published Constraints score comes from Gemini Robotics-ER 1.5 fine-tuned on 200 pairs made 'using the same synthetic data generation recipe and human annotation process'.
A results figure changed between paper versions and the text was not updated
Video scores changed between two versions of the paper. The text was not updated to match.18219+1
Details
In the ASIMOV-2.0 paper of 2025-09-25, Figure 5(a) gives video risk accuracy of 0.89 (Gemini 2.5 Pro), 0.64 (Claude Opus 4.1) and 0.52 (GPT-5). The 2025-11-21 version shows 0.84, 0.64 and 0.44. Both versions say the text-to-video gap for GPT-5 is 40%, which matches the old figure (0.92 to 0.52) but not the new one (0.92 to 0.44, a 48-point gap). Figure 10 also changed; the other result figures are byte-identical. No changelog explains the change.
Many test items have no human labels
Half of the v1 test items have no human labels. No figure for agreement between raters is reported.1212
Details
In v1, 1,140 of the 2,273 validation instructions have human labels, and Multimodal-Manual has none (its data was written by people). The paper describes 'a round of human voting' without voter counts or agreement figures. ASIMOV-2.0 used 5 raters per item and dropped low-consensus items, which also removes the hardest or most disputed cases, and reports no agreement statistic.
Most of the advertised data was never released
Most of the 2.9 million v1 instructions are not public.22116+2
Details
The CoRL abstract describes 500k situations and 3M instructions. The public bucket holds the validation sets, the Dilemmas-SciFi training set and a 7-item RoboPAIR set. The training sets for Multimodal-Auto (288,421 instructions), Injury (2,335,361) and Dilemmas-Auto (262,621) are not in it. A user asked for the Multimodal-Auto training set on 2025-06-22 (issue #2); there is no reply. The paper also mentions a test split for each component, which is not described or released.
Released file counts differ from the paper
Some public v1 files hold fewer items than the paper lists.11716+1
Details
v1 Table 1 lists 319 Injury validation instructions; the released injury_val file holds 304 records with one instruction each. Dilemmas-Auto validation is listed as 100 contexts and 200 instructions; the file holds 34 records. Dilemmas-SciFi training is listed as 9,056 contexts; the file holds 9,004. Some of the gaps may come from different counting units, which the release does not document.
Details
About
- What it is
- Benchmark12
More
Both papers call it a benchmark. v1 also describes large training sets for building robot constitutions (see kind_secondary).
- Built by
- Google DeepMind, Princeton University123
More
v1 authors: Sermanet, Majumdar, Irpan, Kalashnikov, Sindhwani. 2.0 authors: Jindal, Kalashnikov, Hofer, Chang, Garikapati, Majumdar, Sermanet, Sindhwani. All list Google DeepMind (Robotics); Anirudha Majumdar also lists Princeton University.
Google DeepMind · All authors of both papers.12
Princeton University · Second affiliation of Anirudha Majumdar.23
- Released
- March 2025, at CoRL 2025162622+1
More
v1 data uploaded 2025-03-06; arXiv v1 on 2025-03-11. Published at CoRL 2025 (PMLR volume 305, pages 4767 to 4823).
TFDS registered the v1 datasets on 2025-03-07 (tensorflow/datasets commit 'Release of ASIMOV datasets').
- Version
- v1 (March 2025) and 2.0 (September 2025)2831+2
More
Two versions with different questions. v1 (March 2025): Multimodal, Injury and Dilemmas sets. ASIMOV-2.0 (September 2025): Injury (text), Video and Constraints (images). Each released dataset is TFDS version 0.1.0.
The project root redirects to /v1/ or /v2/. Both versions stay downloadable. ASIMOV-Agentic (Google DeepMind, 2026-07) extends the family and has its own Atlas entry.
v1 (ASIMOV-1.0) · Five subsets: Multimodal-Auto, Multimodal-Manual, Injury, Dilemmas-Auto, Dilemmas-SciFi. Plus a 7-item RoboPAIR validation set.116
ASIMOV-2.0 · Three components: Injury (319 text scenarios), Video (287 generated videos), Constraints (164 image and constraint pairs).2
- Last update
- November 2025. The paper was revised.2916
More
ASIMOV-2.0 paper v2 on 2025-11-21. The 2.0 data was uploaded 2025-09-24 and has not changed since.
The 2025-11 revision replaced one results figure (see issues.i3). ASIMOV-2.0 has no conference venue that we found; its arXiv record carries report number GDM-01-01.
- Status
- No updates since November 2025. Still cited. Inferred16298+3
More
No data change since September 2025. Google still cited it in July 2026 and released an agentic extension.
Data unchanged since 2025-09-24; last paper revision 2025-11-21. The code repository has had no commit since 2025-03-11, and its two issues (2025-03-12, 2025-06-22) have no reply. The Gemini Robotics ER 2 model card (2026-07-30) links its semantic-safety claim to the ASIMOV v2 page.
Setup
- Runs in
- Recorded data12
More
Models answer questions about text, images or videos. No robot or simulator runs.
- Robot
- None. The test uses questions only.12
More
No robot under test. Some prompts tell the model it is a robot, for example 'I am a humanoid robot'.
v1 images start from RoboVQA frames taken by single-arm mobile robots and by people wearing head cameras, then edited with Imagen 3.
- Setting
- Mixed Inferred12
More
Injury scenarios come from US emergency-room reports (NEISS 2023). Constraint images show objects with physical properties such as heat, weight or fragility.
- Size
- 2,273 test items in v1 and 770 in v21212
More
v1 validation: 310 contexts or images and 2,273 instructions, 1,140 with human labels. v1 in all: 513,679 contexts and 2,942,060 instructions. 2.0: 319 + 287 + 164 = 770 evaluation items.
v1 Table 1 · Validation: Multimodal-Auto 1,311 instructions (789 human labels), Multimodal-Manual 159 (0), Injury 319 (163), Dilemmas-Auto 200 (35), Dilemmas-SciFi 284 (153).121
v1 released files · Public record counts: injury_val 304, dilemmas_auto_val 34, dilemmas_scifi_val 51, dilemmas_scifi_train 9,004, multimodal_auto_val 50, multimodal_manual_val 59, multimodal_robopair_val 7.1617
ASIMOV-2.0 · Injury 319 text scenarios; Video 287 videos (193 without a realistic injury, 94 with one); Constraints 164 image and constraint pairs in 8 categories.231
Scoring and access
- Scored by
- Accuracy12
More
Agreement with human labels. 2.0 also measures timing error in seconds and a constraint violation rate, which have no taxonomy value.
- Score
- Share of answers that match human raters127
More
v1: 'alignment rate', the share of yes-or-no desirability answers that match human labels, in a normal mode and in an 'adversary' mode where the model is told to flip good and bad. 2.0 Injury: accuracy on four multiple-choice questions (risk type, risk severity, effect of the action, severity after the action). 2.0 Video: risk accuracy, plus error in seconds for the last moment an intervention could prevent injury. 2.0 Constraints: share of answers that point at an object breaking the stated robot limit.
v1's headline 84.3% is the average of normal-mode and adversary-mode alignment. 2.0 labels: 5 raters per item; low-consensus items dropped (Injury), at least 60% consensus and a timestamp spread under 1.0 s (Video), at least 80% consensus (Constraints). No agreement statistic between raters is reported.
- Trials
- One answer per item Inferred129
More
Each model answers each item once; the papers do not report repeated runs.
Not stated explicitly. VLESA (2026-06) applied the official consensus rule and kept 189 of the 287 videos with valid intervention labels for timing metrics.
- Who runs it
- Each team tests its own model Inferred283
More
Data and prompts are public; each team scores its own models.
- Error bars
- Not reported Inferred127
More
Results are single numbers or plotted points without error bars. The only intervals in the 2.0 paper are for a pointing test in its appendix.
- Leaderboard
- None. Scores are only in papers.283
More
No results table on the v1 or v2 project pages or in the repository.
- Code licence
- Unknown
More
asimov-benchmark/code holds a 6-byte README and one notebook, with no LICENSE file; GitHub reports no licence. The TFDS loader scripts in tensorflow/datasets are Apache-2.0, which covers those scripts only.
- Data licence
- Unknown
More
No licence on the v1 or v2 project pages, in the repository, in any dataset_info.json (no licence field) or in the TFDS builder code. The CC BY 4.0 notice on TFDS catalog pages covers the page content. Sources inside the data (NEISS reports, RoboVQA frames, Imagen and Veo outputs) were not checked for terms.
- Asset licence
- Unknown
More
Images and videos are AI-generated (Imagen 3, Veo 3) or edited from RoboVQA frames. No separate licence stated; looked in the same places as license_data.
Sources 31
- 1Generating Robot Constitutions & Benchmarks for Semantic Safety (ASIMOV v1, full text)Paper · Mar 2025 · checked 10 Oct 2026
- 2Can AI Perceive Physical Danger and Intervene? (ASIMOV-2.0, full text v2)Paper · Nov 2025 · checked 10 Oct 2026
- 3ASIMOV Benchmark v2 project pageOfficial site · Sep 2025 · checked 10 Oct 2026
- 4Semantic Scholar API record for arXiv:2503.08663 (v1)Index · 10 Oct 2026 · checked 10 Oct 2026
- 5Semantic Scholar API record for arXiv:2509.21651 (ASIMOV-2.0)Index · 10 Oct 2026 · checked 10 Oct 2026
- 6Gemini Robotics: Bringing AI into the Physical World (Section 5, Figure 29)Paper · Mar 2025 · checked 10 Oct 2026
- 7Gemini Robotics 1.5 tech report (v3; ASIMOV-2.0 section, Figures 18 and 19)Paper · Oct 2025 · checked 10 Oct 2026
- 8Gemini Robotics ER 2 model card (links semantic safety benchmarks to ASIMOV v2)Official site · Jul 2026 · checked 10 Oct 2026
- 9VLESA: Vision-Language Embodied Safety Agent for Human Activity MonitoringPaper · Jun 2026 · checked 10 Oct 2026
- 10Visual Grounding Safety in Vision-Language ModelsPaper · Oct 2026 · checked 10 Oct 2026
- 11What the Guard Misses, the Robot Executes: Implied Harm in VLA InstructionsPaper · Oct 2026 · checked 10 Oct 2026
- 12Gemini Robotics 2: Safety Evaluations (ASIMOV-Agentic report)Paper · Jul 2026 · checked 10 Oct 2026
- 13SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models (ACL 2026 Findings)Paper · Apr 2026 · checked 10 Oct 2026
- 14HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household ScenariosPaper · Mar 2026 · checked 10 Oct 2026
- 15Probing Collision Grounding in Vision-Language Models for Safe Human-Robot CollaborationPaper · May 2026 · checked 10 Oct 2026
- 16Google Cloud Storage listing, gresearch bucket, prefix robotics/asimov (object names, sizes, upload dates)Repository · 24 Sep 2025 · checked 10 Oct 2026
- 17asimov_injury_val dataset_info.json (record counts; other datasets read at the same path pattern)Repository · Mar 2025 · checked 10 Oct 2026
- 18ASIMOV-2.0 full text, version 1Paper · Sep 2025 · checked 10 Oct 2026
- 19ASIMOV-2.0 v1 Figure 5(a): risk recognition, text vs videoPaper · Sep 2025 · checked 10 Oct 2026
- 20ASIMOV-2.0 v2 Figure 5(a): risk recognition, text vs videoPaper · Nov 2025 · checked 10 Oct 2026
- 21ASIMOV v1, CoRL 2025 camera-ready PDFPaper · Sep 2025 · checked 10 Oct 2026
- 22ASIMOV v1, CoRL 2025 proceedings page (PMLR 305)Paper · Sep 2025 · checked 10 Oct 2026
- 23ASIMOV_Datasets.ipynb (official loading notebook and dataset list)Repository · Mar 2025 · checked 10 Oct 2026
- 24asimov-benchmark/code issue #2: Multimodal-Auto train release (no reply)Repository · Jun 2025 · checked 10 Oct 2026
- 25TensorFlow Datasets catalog: asimov_injury_valRepository · 2026 · checked 10 Oct 2026
- 26ASIMOV v1 arXiv abstract page (submission history)Paper · Mar 2025 · checked 10 Oct 2026
- 27TFDS builder source for ASIMOV and ASIMOV v2 datasetsRepository · 24 Sep 2025 · checked 10 Oct 2026
- 28ASIMOV Benchmark v1 project pageOfficial site · Mar 2025 · checked 10 Oct 2026
- 29ASIMOV-2.0 arXiv abstract page (v1 2025-09-25, v2 2025-11-21)Paper · Sep 2025 · checked 10 Oct 2026
- 30GitHub API: asimov-benchmark/code (stars, forks, licence, last push)Index · 10 Oct 2026 · checked 10 Oct 2026
- 31asimov_v2_videos dataset_info.json (287 records)Repository · 24 Sep 2025 · checked 10 Oct 2026
Where we searched for missing information
validity and sim_to_real: v1 arXiv and PMLR camera-ready; ASIMOV-2.0 arXiv v1 and v2 (all figures viewed); project pages v1 and v2; Gemini Robotics (2503.20020), Gemini Robotics 1.5 (2510.03342), Gemini Robotics ER 2 model card and release post, Gemini Robotics 2 safety report; all 43 and 15 Semantic Scholar citing papers scanned by title and 20 opened (VLESA, Visual Grounding Safety, SafetyALFRED, HomeSafe-Bench, Veo world-simulator paper 2512.10675 and others). The scores are themselves agreement with human raters; no study compares ASIMOV scores with robot behaviour, and no rater-agreement statistic is published. Validity list left empty.
license_code, license_data, license_assets: asimov-benchmark/code tree and licence field; v1 and v2 project pages; dataset_info.json for all 11 released datasets; TFDS catalog pages; TFDS builder source (asimov.py, asimov_v2.py).
issues.i1 (training sets): Public bucket listing for prefix robotics/asimov (55 objects, 11 datasets), the official notebook's dataset list, GitHub issue #2.
ASIMOV-2.0 venue: arXiv record (report number GDM-01-01, no journal reference), v2 project page, web search. No conference or journal found.
Gemini Robotics 2 / ER 2 ASIMOV-2.0 numbers: ER 2 model card (HTML and PDF), Gemini Robotics 2 release post, Gemini Robotics 2 safety report. Only ASIMOV-Agentic numbers are given; ASIMOV v2 is linked without numbers.
Change history
- Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/frontier-labs.json. New findings: most v1 training data unreleased; ASIMOV-2.0 Figure 5(a) changed between arXiv versions without the text; Google post-trains on ASIMOV-type data; outside uses by VLESA and an Apple co-authored paper. ASIMOV-Agentic kept as a separate entry.
- Published as a full entry.