ASIMOV-Agentic

How to read this picture

Tests whether a robot's planning model refuses unsafe tasks, stops near people and asks for help when unsure.1

Sources
Last checked 10 Oct 2026Basic entry11 of 18 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
Safety
Robot
No body1
Licence
CC-BY-4.0 (inferred)2

Comparisons with real robots

Tried on real robots, not compared1

Details

One real-world check by Google: Apollo 2 humanoid safe stopping (99% detection, 96% safe-pose reliability, lab settings). No correlation statistic between offline scores and real behaviour.

Details

About

What it is
Benchmark Inferred1
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
Google DeepMind (Gemini Robotics Team)1
More

Report 'Gemini Robotics 2: Safety Evaluations', Gemini Robotics Team, Google DeepMind; contributors listed alphabetically. Dataset under the 'google' organisation on Hugging Face.

Released
2026-07 (HF dataset created 2026-07-21; safety report dated 2026-07-29; announced in GR 2 blog 2026-07-30)2
Version
Initial release (no version label) Inferred2
More

Report says future versions will test longer contexts and 'attention jailbreaking'.

Last update
2026-07 (HF dataset last modified 2026-07-24)2

Setup

Runs in
Recorded data1
More

'offline one-step and multiturn (using VLA tool emulation) benchmark'. A Gemini-based VLA confidence emulator replaces real robots in the multi-turn variant.

Robot
No body1
Setting
Tabletop, Industrial, Mixed Inferred1
More

ALOHA tabletop scenes, industrial instrument reading (gauges, sight glasses), garage sorting task.

Size
Six components; exact item counts not stated in the report. Hugging Face size category tag: n<1K. Inferred2
More

Files include constraints/physical_constraints.parquet, gauge_reading/*.parquet and ~30 human_safety_monitoring/*.parquet episodes. The dataset card text is behind a Hugging Face login, so per-file counts were not read.

Scoring and access

Scored by
Accuracy1
More

Self-reported. Frontier models compared: Gemini Robotics ER 2, Claude Opus 4.8, GPT 5.5. Agentic human-monitoring: no model achieves both low FNR and low FPR; FPR under 5% comes with FNR above 40%.

Leaderboard
None. Scores are only in papers.1
More

Results only in the safety report figures.

Code licence
CC-BY-4.0 (inferred) Inferred2
More

Evaluation scripts (asimov_agentic_evals.py etc.) ship inside the same HF dataset repo, whose card declares CC-BY-4.0; no separate code licence seen.

Data licence
CC-BY-4.02
More

HF card metadata license: cc-by-4.0. Access gated ('auto' approval) on Hugging Face.

Access
Download after registering2
More

Gated with automatic approval; requires a Hugging Face login.

Sources 2

  1. 1https://storage.googleapis.com/deepmind-media/gemini-robotics/Gemini-Robotics-2-Safety.pdfOfficial site · checked 10 Oct 2026
  2. 2google/asimov_agentic on Hugging Face (dataset)Index · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).