SIMA 2 evaluation
SIMA 2 evaluation (SIMA Evaluation Suite 2.0)
Google DeepMind's internal test of a Gemini-based game agent: task success in 3D video games, compared with human players.1
Comparisons with real robots
No comparison found Unknown
Details
No physical robot in the loop and no transfer study found (paper, blog). Real-world validity in the robotics sense does not apply directly to a game agent.
Details
About
- What it is
- In-house test Inferred2
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- Google DeepMind (SIMA Team)3
More
Title block: 'SIMA Team, Google DeepMind'. arXiv lists Adrian Bolton and 64 other authors.
- Version
- SIMA Evaluation Suite 2.0 (as described in arXiv 2512.04797v1)1
More
Section 3.4 names 'SIMA Evaluation Suite 2.0'. Changes from SIMA 1: more domains, more tasks ('often by an order of magnitude' for programmatic evals), success text must persist several seconds, limits on actions after completion, sequential chains must be fully completed.
- Last update
- 2025-12: arXiv 2512.04797 v1 (no later version); tech report PDF dated 2025-12-052
More
arXiv lists only v1. The PDF linked from the blog carries the date 2025-12-05.
Setup
- Runs in
- Simulation Inferred1
More
Closed-loop keyboard-and-mouse control inside game engines and research environments. These are game engines, not physics simulators built for robotics (see taxonomy_friction).
- Robot
- Game character1
- Setting
- Game world1
More
Training: Construction Lab, Playhouse, WorldLab (research environments); Goat Simulator 3, Hydroneer, No Man's Sky, Satisfactory, Space Engineers, Valheim, Wobbly Life (commercial). Held-out: ASKA, MineDojo (Minecraft); The Gunk and Genie 3 worlds only qualitatively.
- Size
- 10 training domains (3 research environments + 7 commercial games); held-out quantitative evaluation on ASKA and a 50-task MineDojo subset (15 random seeds per task). Total number of evaluation tasks not stated.1
More
Blog lists more partner games (also Eco, The Gunk, SteamWorld Build, Road 96, Teardown) than the paper's 7 commercial training games; the paper's list is used for the evaluation.
Scoring and access
- Scored by
- Success rate1
More
Organiser-run, self-reported by Google DeepMind. Human baselines are shown with and without the agent's time limit. Uncertainty statistics are not described in the text. Values are the bar labels printed in Fig. 6 of the PDF. Bar-to-series mapping read from bar order; consistent with the text's 'SIMA 2 effectively doubles the average success rate of SIMA 1'. Baseline Gemini models without SIMA training: Flash-Lite 3.2%, Pro 7.0% over the 10 training domains (text, https://arxiv.org/html/2512.04797). Human figures are for 'a representative subset of tasks'.
- Leaderboard
- None. Scores are only in papers.1
More
Results appear only in the paper and blog. No public leaderboard found.
- Code licence
- not released1
More
No evaluation code, task list or repository mentioned in the paper or blog. arXiv's CC BY 4.0 licence covers the paper text only.
- Data licence
- not released1
More
No evaluation tasks or data released. Commercial games are third-party products.
- Access
- Closed1
More
SIMA 2 was released as a 'limited research preview' to a small cohort of academics and game developers. Games used under licensed agreements with their developers.
Sources 5
- 1SIMA 2: A Generalist Embodied Agent for Virtual Worlds (full text)Paper · Dec 2025 · checked 10 Oct 2026
- 2SIMA 2: A Generalist Embodied Agent for Virtual WorldsPaper · Dec 2025 · checked 10 Oct 2026
- 3https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/SIMA_Tech_Report_2025.pdfOfficial site · checked 10 Oct 2026
- 4SIMA 2: A Gemini-Powered AI Agent for 3D Virtual Worlds — Google DeepMindOfficial blog · checked 10 Oct 2026
- 5Scaling Instructable Agents Across Many Simulated WorldsPaper · Mar 2024 · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).