Introduction · about 10 minutes

Basics

An introduction for readers who are new to robotics and AI.

01

What embodied AI is

Embodied AI is AI that controls a physical body, such as a robot arm, a mobile robot or a humanoid robot.

  • It sees the world through cameras and other sensors.
  • It decides what to do next. The part that makes this decision is called the policy.
  • It acts through motors that move arms, hands, wheels or legs.
  • It then sees the result of its action and decides again.
SeeCAMERAS DecidePOLICY ActMOTORS the world
02

What a benchmark is

A benchmark is a fixed set of tasks with a fixed way of scoring them, so that different systems can be compared.

10 TRIES · 8 WORKED 80% 80%Policy A60%Policy B35%Policy C
  • Each task has a goal, for example putting a bowl on a plate.
  • The robot or model tries each task many times.
  • The score is usually the share of attempts that succeeded. This is called the success rate.
  • For example, LIBERO has 130 tasks.
03

Where a test can run

Tests run in one of five settings. They range from no robot at all to real robots.

Recorded data

No robot moves. The system answers questions about recorded video or predicts what happens next.

18 in this atlas

Simulation

A virtual robot carries out the tasks in a scene computed by a physics simulator.

44 in this atlas

AI simulator

A trained neural network predicts what the robot would see after each action. It takes the place of a physics simulator.

2 in this atlas

Sim + real

The benchmark has both simulated tests and real-robot tests.

3 in this atlas

Real robots

Physical robots carry out the tasks.

12 in this atlas
04

Why simulation results may not hold on real robots

A simulator is cheaper and easier to repeat than a real test. It is never an exact copy of the real world.

SIMULATION REALITY
  • Physics in a simulator is approximate, especially friction, contact and soft objects.
  • Simulated images differ from camera images in light, texture and noise.
  • Simulated rooms are usually tidier than real homes and workplaces.
  • Because of these differences, a policy that does well in simulation can do worse on a real robot. This difference is called the sim-to-real gap.
05

How a simulator is checked against real robots

Researchers run the same policies in the simulator and on real robots, then compare the two sets of scores.

  • First, several policies are scored in the simulator.
  • Next, the same policies are scored on real robots.
  • If the policies rank in the same order in both, the simulator is a useful stand-in for real tests. If the order is different, it is not.
Illustration, not real data SIM SCORE → REAL SCORE ↑
On the benchmark mapReal robots: 15Checked against real robots: 8Not checked: 56

For example, LIBERO has been checked against real robots once. The study measured a correlation of Correlation r = 0.66 to 0.70 across 5 policies. Its authors described the link as poor.

A correlation of 1 means the two sets of scores agree perfectly. A correlation of 0 means they are unrelated.

06

Common ways to test robot AI

Labs and companies use five main approaches.

Simulation benchmarks

A fixed set of tasks in a physics simulator. Anyone can download the benchmark and run it.

Needs a robot: NoWho runs the test: You
40in this atlas

Simulators checked against real robots

Simulation benchmarks whose scores have been compared with real-robot results for the same policies.

Needs a robot: Only for the comparisonWho runs the test: You
8in this atlas

Real-robot evaluation services

You send your policy to an organiser. The organiser runs it on their robots and publishes the result.

Needs a robot: The organiser’sWho runs the test: The organiser
3in this atlas

AI simulators

A trained world model predicts what the robot would see after each action. The policy is scored inside that prediction.

Needs a robot: NoWho runs the test: You or the builder
2in this atlas

Recorded-data tests

The system answers questions about recorded video, or predicts future video. No robot moves.

Needs a robot: NoWho runs the test: You
18in this atlas

There are also competitions, such as the BEHAVIOR Challenge, and one national standard, YD/T 6770-2026.

Opinion

In general, the closer a test is to real robots, the slower and more expensive it is, and the more its scores tell you about real-world performance.

This is our reading of the five approaches above.
07

What to check before trusting a score

These four questions apply to any benchmark result.

Are scores near 100%?

When most systems score close to the maximum, small differences between them are not meaningful.

1050

Was the same protocol used?

Papers often run the same benchmark with different numbers of trials, training data or settings.

Who ran the test?

Results run by the organisers are easier to compare than results each team reports on its own.

Are error bars reported?

Without a measure of uncertainty, you cannot tell a real improvement from chance.

All four questions apply to LIBERO. Its page shows the details.

08

Glossary

Benchmark
A fixed set of tasks, rules and a score, so systems can be compared.
Closed-loop
The policy acts, sees the result, acts again. Errors add up, as on a real robot.
Correlation (r)
How closely two sets of scores move together. 0 means no link, 1 means a perfect one.
Demonstration
A recorded example of a task being done, used for training.
Embodied AI
AI that controls a body in the physical world, such as a robot.
Embodiment
The robot body: its arms, hands, base and sensors.
Episode
One try at a task, from start to success, failure or time-out.
Error bars
A range showing how much a score could move by chance.
Leaderboard
A public table of results.
Lifelong learning
Learning tasks one after another without forgetting earlier ones.
Locomotion
Walking, balance and whole-body movement.
Manipulation
Handling objects: grasping, moving, placing, using.
Policy
The AI that picks a robot’s next move from what it sees.
Protocol
The exact way a test is run: trials, seeds, training data.
Real-to-sim
A simulation built to copy a specific real setup.
Saturation
Top systems all score near the maximum, so the test stops telling them apart.
Self-reported
Each team runs the test and publishes its own number.
Sim-to-real gap
The difference between how a policy does in simulation and on a real robot.
Simulation
Software that computes physics, so a virtual robot can act in a virtual scene.
Success rate
Tries that worked, divided by all tries.
Teleoperation
A person drives the robot remotely to record examples.
VLA
Vision-language-action model. Takes camera images and a text instruction, outputs robot actions.
World model
A neural network that predicts what happens next, often as video.
09

How to read this site

3 Source number
Hover over the number to see the source. Click it to go to the list of sources.
Inferred
We worked this out from the sources. The details explain how.
Secondary
Only reported by others, such as news articles.
Unknown
We searched the primary sources and did not find this information.
Opinion
Our own judgement. It is kept separate from the facts.

A fact without a label was checked against a primary source. How we check facts

Pictures

Each benchmark has a picture drawn from its own facts. The same parts are used for every benchmark, so pictures can be compared.

Example: LIBERO
Background: where the test runs. Blue grid: simulation. Blue dots: AI simulator. Split: simulation and real robots. Green: real robots. Film frames: recorded data.
Figure
Figure: the robot. For example one arm, two arms, a humanoid, a four-legged robot or a camera when no robot body is used.
Objects
Objects: the setting. A table, a kitchen, a living room, shelves, a factory line or a game world.
Extras
Extras: what is tested. A dashed path for navigation, a person for working with people, a question for reasoning, a marked area for safety, stacked frames for world models.
Corner mark: checked against real robots. The same three marks as on the benchmark map.