EmbodiedGovBench
Scores embodied agent systems on governance: permission limits, recovery, upgrades, human override and audit trails, in AI2-THOR scenarios.1
Last checked 10 Oct 2026Basic entry11 of 14 facts checked at the sourceNext check 8 Apr 2027
Comparisons with real robots
Not checked3
Details
Authors state external validity (prediction of real deployment governance) is undemonstrated.
Details
About
- What it is
- Benchmark Inferred1
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- Harbin Institute of Technology (3 authors), Heriot-Watt University Malaysia campus, Soochow University.3
More
Repo README citation adds a sixth author (Zeyd Boukhers) not on arXiv v1.
- Released
- 2026-04 (arXiv v1 2026-04-13).1
- Last update
- 2026-05-05: single-commit code release 'EmbodiedGovBench v0.1 (paper-submission artifact bundle)', tag v0.1.0-jair-submission.2
More
Tag name suggests a JAIR submission; acceptance not found.
Setup
Scoring and access
Sources 3
- 1EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent SystemsPaper · Apr 2026 · checked 10 Oct 2026
- 2s20sc/embodied-gov-bench on GitHub (repository)Repository · checked 10 Oct 2026
- 3EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems (full text)Paper · Apr 2026 · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).