About

About this site

How the information on this site is collected, checked and labelled.

What this site is

This site is a guide to the benchmarks used to test AI that controls robots. It has 79 entries. 30 of them are full entries, which also cover known problems, comparisons with real robots and our assessment. The other entries are basic entries with the main facts.

Sienna Chen built the site. The site is based on research done with VLGE in 2026.

Sources

We use primary sources only. These are the original paper, the official website, the official code or data, and the official leaderboard or blog. We use news articles and other people’s surveys only to find primary sources. We do not use them as evidence.

Each fact links to its source and shows the date we checked it.

Labels on facts

3 Number
We checked this fact at the source. Hover over the number to see the source.
Inferred
We worked this out from the sources. The details explain how.
Secondary
Only other people report this. We did not find it in a primary source.
Unknown
We searched and found no information. We do not guess.
Opinion
This is our assessment, with our reasons. It is kept separate from the facts.
Disputed
Sources disagree. Both views are shown.

Checked against real robots

We mark a simulation benchmark as checked only when someone scored the same policies in the benchmark and on real robots, and then compared the two sets of scores. We show the result of each comparison with its number.

How often entries are checked

Each entry shows when it was last checked and when the next check is due. Entries are checked again every 180 days.

Each week, a script lists new papers and changes to the sources. A person reviews this list before anything on the site changes. Nothing is published without review. Past changes are listed in the changelog.

Report an error

If you find an error, send us an email with a link to the source.