Introduction · about 10 minutes
Basics
An introduction for readers who are new to robotics and AI.
What embodied AI is
Embodied AI is AI that controls a physical body, such as a robot arm, a mobile robot or a humanoid robot.
- It sees the world through cameras and other sensors.
- It decides what to do next. The part that makes this decision is called the policy.
- It acts through motors that move arms, hands, wheels or legs.
- It then sees the result of its action and decides again.
What a benchmark is
A benchmark is a fixed set of tasks with a fixed way of scoring them, so that different systems can be compared.
- Each task has a goal, for example putting a bowl on a plate.
- The robot or model tries each task many times.
- The score is usually the share of attempts that succeeded. This is called the success rate.
- For example, LIBERO has 130 tasks.
Where a test can run
Tests run in one of five settings. They range from no robot at all to real robots.
Recorded data
No robot moves. The system answers questions about recorded video or predicts what happens next.
18 in this atlasSimulation
A virtual robot carries out the tasks in a scene computed by a physics simulator.
44 in this atlasAI simulator
A trained neural network predicts what the robot would see after each action. It takes the place of a physics simulator.
2 in this atlasSim + real
The benchmark has both simulated tests and real-robot tests.
3 in this atlasWhy simulation results may not hold on real robots
A simulator is cheaper and easier to repeat than a real test. It is never an exact copy of the real world.
- Physics in a simulator is approximate, especially friction, contact and soft objects.
- Simulated images differ from camera images in light, texture and noise.
- Simulated rooms are usually tidier than real homes and workplaces.
- Because of these differences, a policy that does well in simulation can do worse on a real robot. This difference is called the sim-to-real gap.
How a simulator is checked against real robots
Researchers run the same policies in the simulator and on real robots, then compare the two sets of scores.
- First, several policies are scored in the simulator.
- Next, the same policies are scored on real robots.
- If the policies rank in the same order in both, the simulator is a useful stand-in for real tests. If the order is different, it is not.
For example, LIBERO has been checked against real robots once. The study measured a correlation of Correlation r = 0.66 to 0.70 across 5 policies. Its authors described the link as poor.
A correlation of 1 means the two sets of scores agree perfectly. A correlation of 0 means they are unrelated.
Common ways to test robot AI
Labs and companies use five main approaches.
Simulation benchmarks
A fixed set of tasks in a physics simulator. Anyone can download the benchmark and run it.
Simulators checked against real robots
Simulation benchmarks whose scores have been compared with real-robot results for the same policies.
Real-robot evaluation services
You send your policy to an organiser. The organiser runs it on their robots and publishes the result.
AI simulators
A trained world model predicts what the robot would see after each action. The policy is scored inside that prediction.
There are also competitions, such as the BEHAVIOR Challenge, and one national standard, YD/T 6770-2026.
In general, the closer a test is to real robots, the slower and more expensive it is, and the more its scores tell you about real-world performance.
What to check before trusting a score
These four questions apply to any benchmark result.
Are scores near 100%?
When most systems score close to the maximum, small differences between them are not meaningful.
Was the same protocol used?
Papers often run the same benchmark with different numbers of trials, training data or settings.
Who ran the test?
Results run by the organisers are easier to compare than results each team reports on its own.
Are error bars reported?
Without a measure of uncertainty, you cannot tell a real improvement from chance.
All four questions apply to LIBERO. Its page shows the details.
Glossary
- Benchmark
- A fixed set of tasks, rules and a score, so systems can be compared.
- Closed-loop
- The policy acts, sees the result, acts again. Errors add up, as on a real robot.
- Correlation (r)
- How closely two sets of scores move together. 0 means no link, 1 means a perfect one.
- Demonstration
- A recorded example of a task being done, used for training.
- Embodied AI
- AI that controls a body in the physical world, such as a robot.
- Embodiment
- The robot body: its arms, hands, base and sensors.
- Episode
- One try at a task, from start to success, failure or time-out.
- Error bars
- A range showing how much a score could move by chance.
- Leaderboard
- A public table of results.
- Lifelong learning
- Learning tasks one after another without forgetting earlier ones.
- Locomotion
- Walking, balance and whole-body movement.
- Manipulation
- Handling objects: grasping, moving, placing, using.
- Policy
- The AI that picks a robot’s next move from what it sees.
- Protocol
- The exact way a test is run: trials, seeds, training data.
- Real-to-sim
- A simulation built to copy a specific real setup.
- Saturation
- Top systems all score near the maximum, so the test stops telling them apart.
- Self-reported
- Each team runs the test and publishes its own number.
- Sim-to-real gap
- The difference between how a policy does in simulation and on a real robot.
- Simulation
- Software that computes physics, so a virtual robot can act in a virtual scene.
- Success rate
- Tries that worked, divided by all tries.
- Teleoperation
- A person drives the robot remotely to record examples.
- VLA
- Vision-language-action model. Takes camera images and a text instruction, outputs robot actions.
- World model
- A neural network that predicts what happens next, often as video.
How to read this site
- 3 Source number
- Hover over the number to see the source. Click it to go to the list of sources.
- Inferred
- We worked this out from the sources. The details explain how.
- Secondary
- Only reported by others, such as news articles.
- Unknown
- We searched the primary sources and did not find this information.
- Opinion
- Our own judgement. It is kept separate from the facts.
A fact without a label was checked against a primary source. How we check facts
Pictures
Each benchmark has a picture drawn from its own facts. The same parts are used for every benchmark, so pictures can be compared.
- Background: where the test runs. Blue grid: simulation. Blue dots: AI simulator. Split: simulation and real robots. Green: real robots. Film frames: recorded data.
- Figure
- Figure: the robot. For example one arm, two arms, a humanoid, a four-legged robot or a camera when no robot body is used.
- Objects
- Objects: the setting. A table, a kitchen, a living room, shelves, a factory line or a game world.
- Extras
- Extras: what is tested. A dashed path for navigation, a person for working with people, a question for reasoning, a marked area for safety, stacked frames for world models.
- Corner mark: checked against real robots. The same three marks as on the benchmark map.