WAGIBench

How to read this picture

Tests whether vision-language models can guess a smart-glasses wearer's goal from video, audio and phone context.1

Sources
Last checked 10 Oct 2026Basic entry13 of 20 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data2
Checked against real robots
Not checked
Skill
Reasoning
Robot
No body2
Licence
Apache-2.03

Comparisons with real robots

No comparison found Unknown

Details

Not applicable in practice: no robot, no physical action.

Details

About

What it is
Benchmark Inferred1
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
Meta Reality Labs; Meta FAIR2
More

Each author is marked Meta Reality Labs or Meta FAIR (a few authors show no affiliation).

Released
2025-10 (arXiv v1 2025-10-25; repo initial commit 2025-10-22; dataset tarball last-modified 2025-10-23)1
More

Repo created 2025-07-02 per GitHub API (https://api.github.com/repos/facebookresearch/WAGIBench), first commit 2025-10-22.

Version
OB2 v2.4.4 data files (e.g. ob2_v2.4.4_1_mcq.json)4
Last update
2026-02: last commit 2026-02-12 (dependency version bumps); data access instructions updated 2025-10-275
More

No dataset version change found.

Setup

Runs in
Recorded data2
More

Models answer from recorded observations; no control loop.

Robot
No body2
Setting
Mixed Inferred2
More

Scripted everyday activities (memory, learning, health and fitness, meal preparation, chores, recreation).

Size
Unknown62
More

arXiv v1: 29 hours, 348 participants, 3,477 recordings. NeurIPS version: ~30 hours, 363 participants, 3,482 recordings.6

~7k multiple-choice questions (one 'similar' and one 'dissimilar' MCQ per sample, 3 distractors each); human study subset 586 samples; dataset tarball 17,085,909,384 bytes2

Scoring and access

Scored by
Accuracy, Automatic judge2
More

Paper reports bootstrapped 95% confidence intervals of the mean (Fig. 4). Self-reported. Repo reference run: Qwen2.5-VL-72B MCQ 0.871, generative 0.507 (https://github.com/facebookresearch/WAGIBench/blob/main/README.md).

Leaderboard
None. Scores are only in papers.4
More

Repo gives raw predictions for paper models; no leaderboard.

Code licence
Apache-2.03
More

LICENSE file: Apache License 2.0.

Data licence
Unknown
More

README says 'This project is licensed under the Apache 2.0 license' but names no separate data licence; the 17 GB tarball was not downloaded to look for one. Paper says participants consented to public release.

Access
Open download4
More

Direct download link, no gating seen.

Published at
NeurIPS 2025 Datasets and Benchmarks Track (spotlight per arXiv comment)6

Sources 6

  1. 1Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable AgentsPaper · Oct 2025 · checked 10 Oct 2026
  2. 2Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents (full text)Paper · Oct 2025 · checked 10 Oct 2026
  3. 3facebookresearch/WAGIBench on GitHub (blob)Repository · checked 10 Oct 2026
  4. 4facebookresearch/WAGIBench on GitHub (file README.md)Repository · checked 10 Oct 2026
  5. 5facebookresearch/WAGIBench on GitHub (repository)Index · checked 10 Oct 2026
  6. 6Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable AgentsOfficial site · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).