Meta-World

Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

How to read this picture

Meta-World is a set of 50 simulated tabletop tasks for one Sawyer robot arm. It is used to test multi-task and meta reinforcement learning (learning to adapt quickly to new tasks), and recent vision-language-action (VLA) papers also report results on it.123

Sources
Last checked 10 Oct 2026Full entry53 of 66 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm1
Sawyer (simulated)
Used by
RL papers and small VLA models245+3
1,843 citations
Licence
MIT910
Commercial use: allowed

What a score here does not tell you Inferred

  1. How well a policy (the robot's control model) will do on a real robot.We found no study that compares Meta-World scores with real-robot scores for the same policies.
  2. How well a policy copes with visual changes.All tasks use one fixed scene. Policies are usually tested on goal positions seen in training.
  3. Whether results from different papers can be compared.Papers use different reward versions and different rules for averaging scores.
ChartPublished scores over time
95% AND ABOVE010203040506070809010020252026Diffusion Policy · 10.5% · 2024-09TinyVLA-H · 31.6% · 2024-09π0 (run by SmolVLA team) · 47.9% · 2025-06SmolVLA (2.25B) · 68.24% · 2025-06πRL on π0 (Flow-Noise) · 85.8% · 2025-10Evo-1 · 80.6% · 2025-11AnoleVLA · 67.85% · 2026-03Evo-Depth · 84.4% · 2026-05LA4VLA · 87.53% · 2026-06FabriVLA · 90% · 2026-07Diffusion Policy 10.5%FabriVLA 90%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.11412+5

Comparisons with real robots

Tried on real robots, not compared Inferred124+3

Details

Neither the original paper nor Meta-World+ has real-robot experiments. PolaRiS cites Meta-World among simulation benchmarks that fail to capture real-world visual complexity, without measuring it. The basic entry graded this 'none-found'; we use 'demonstrated' because the taxonomy defines it as 'some policies also ran on real robots'.

Our assessment Opinion

Compare numbers only when they use the same protocol and version.

Reasoning

A 'Meta-World score' can come from multi-task reinforcement learning (RL) with V1 or V2 rewards, from meta-RL on held-out tasks, or from an imitation-learning tier average. These are different measurements. Compare numbers only when they use the same protocol, version and averaging rule.

Confidence: high

A high score shows that a policy can do these tasks in one fixed simulated scene.

Reasoning

A high score shows that a policy can handle 50 short tabletop tasks in one fixed simulated scene, usually on goal positions seen in training. It says little about performance on real robots or about how well the policy copes with visual changes.

Confidence: medium

Meta-World is actively maintained. Papers should state the versions they used so that others can reproduce their results.

Reasoning

Meta-World is actively maintained and has numbered versions. This makes it a reasonable choice for RL research that needs reproducible results, as long as papers state the package version, the reward version and the MuJoCo version.

Confidence: medium

Read the score for each tier as well as the average.

Reasoning

For VLA results, read the score for each difficulty tier. The headline average gives the 5 very hard tasks as much weight as the 28 easy ones. Small gains in the average can come from a few tasks.

Confidence: medium

Known problems 5

  1. Protocol details differ between the paper, the docs and later papers

    Sources disagree on whether goal positions are fixed or random. They also give different numbers of test goals.11920+2

    Details

    The original paper and the benchmark docs say MT10 and MT50 goal positions are fixed (the docs carry a 'TODO: check this'), while the evaluation docs cycle through 50 goal positions per task; MOORE and PaCo report 'MT10-rand' with random goals. ML1 is described with 50 held-out positions (paper), 10 (benchmark docs) and 40 (evaluation docs). The original paper reports the average of the maximum success reached during training. Horizons differ (500 steps in the docs, 400 in FabriVLA).

  2. VLA papers average four difficulty tiers instead of 50 tasks

    VLA papers take the plain mean of four difficulty-tier averages. This gives the 5 very hard tasks as much weight as the 28 easy ones. Inferred41612+2

    Details

    Recent VLA papers report the plain mean of four difficulty-tier averages (28 easy, 11 medium, 6 hard, 5 very hard tasks), so each very hard task weighs about 5.6 times as much as an easy one. For SmolVLA (0.45B) the tier mean is 57.3%, while a task-weighted mean of the same tiers is about 66.8% (our arithmetic). The same model also appears with different numbers: π0 at 47.9, 50.5, 50.8 and 47.91; SmolVLA at 57.3 or 68.2 depending on model size.

  3. Scores on some protocols are approaching 100%

    In VLA papers, the average over the four tiers reaches 90%. Meta-learning scores stay near 30% to 40%.111615+2

    Details

    On the VLA tier-average protocol, reported scores rose from 31.6% (TinyVLA, 2024-09) to 90.0% (FabriVLA, 2026-07), with individual tiers at up to 100%. MT10 reaches 88.7% (MOORE, 2024). Meta-learning remains far from the ceiling: Meta-World+ re-runs with V2 rewards give 25.7% to 36.76% on ML10 and ML45 (MAML and RL2).

  4. Results depend on the MuJoCo version

    Newer versions of MuJoCo (the physics simulator) change how contacts between objects are represented. For this reason, v3.1.1 requires MuJoCo 3.3.0.2310

    Details

    v3.1.1 pins mujoco==3.3.0 because newer MuJoCo versions change how contacts between objects are represented. The same release fixed the button-press dense rewards. Results under other MuJoCo versions, or before the fix, may differ.

  5. Two reward versions made published results incomparable

    The original V1 rewards (the feedback signal used in reinforcement learning) were replaced by V2 rewards without documentation. The two versions give different scores, so older results cannot be compared.221

    Details

    Meta-World+ reports that the original 'V1' dense rewards were overwritten by 'V2' rewards without documentation, so papers used different rewards. In its re-runs every multi-task method scored higher with V2: PaCo reached 26.2% on MT10 with V1 and 73.6% with V2. Published numbers disagree too: PaCo's MT10 is 71.6 in its own paper and 85.4 as reported by MOORE; PCGrad's published 90% on MT10 could not be reproduced (70.5%), likely because it used a Meta-World version no longer available.

Details

About

What it is
Benchmark12
More

A suite of 50 tasks with defined multi-task and meta-learning evaluation protocols.

Built by
Stanford University, UC Berkeley, Columbia University, University of Southern California, Robotics at Google, Farama Foundation (maintainer since 2023)1232
More

Original authors (2019) · Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Avnish Narayan, Hayden Shively, Adithya Bellathur, Karol Hausman, Chelsea Finn, Sergey Levine: Stanford, UC Berkeley, Columbia, USC, Robotics at Google.1

Farama Foundation · Maintains the code since 2023; release v2.0.0 (2023-06-16) is 'the code before the Farama Foundation started maintaining Meta-World'.23

Meta-World+ authors (2025) · Toronto Metropolitan University, University of Surrey, Hamburg University of Technology, Columbia, Google DeepMind, USC, Farama Foundation, Université de Montréal / Mila.2

Released
October 2019, at CoRL 20192425
More

arXiv v1 on 2019-10-24; CoRL 2019. The paper was updated as arXiv v2 on 2021-06-14.

GitHub repository created 2019-09-09.

Version
3.1.1, released June 20262310
More

Environment names now end in '-v3'.

v2.0.0 (2023-06-16) · Code as it was before Farama took over maintenance.23

3.0.0 (2025-06-14) · MuJoCo Python bindings replace mujoco-py; Gymnasium 1.0 API; the original 'V1' reward functions exposed again as an option; V1 environments removed.23

v3.1.1 (2026-06-28) · Pins mujoco==3.3.0 because newer MuJoCo changes how contacts are represented; fixes button-press dense rewards; adds a camera. Release notes say v3.1.0 was skipped, but PyPI lists 3.1.0 (uploaded 2026-06-26).2310

Meta-World+ (paper, 2025) · Describes the v3 changes, adds MT25 and ML25 task sets and custom task sets.2

Last update
October 2026. The last release was v3.1.1 in June 2026.262325
More

Last commit 2026-10-09 (drop Python 3.10). Last release v3.1.1 on 2026-06-28.

Status
Active. There were commits in October 2026. Inferred262325
More

Commits in October 2026; release in June 2026.

Within the six-month rule.

Setup

Runs in
Simulation1
Simulator
MuJoCo12310
More

Originally mujoco-py; official MuJoCo bindings since 3.0.0.

Robot
One arm1
Robot model
Sawyer (simulated)12
More

Actions: 3D end-effector displacement plus a gripper value. Observations: 39-dimensional state including object and goal positions; goal zeroed for meta-learning.

Setting
Tabletop127
Tasks
50 tasks128
More

50 tasks; task sets MT1, MT10, MT50, ML1, ML10, ML45, plus MT25 and ML25 since 2025

The repository has 50 '_v3' task environment files (counted by us).

Scenes
1 scene Inferred128
More

One shared tabletop scene; tasks differ in objects

Training data
There is no fixed dataset. Each task has a scripted expert policy.136+1
More

No official dataset. Tasks have scripted expert policies; imitation-learning papers generate their own demonstrations (often 50 per task).

lerobot/metaworld_mt50 · Hugging Face dataset (Apache-2.0 tag) used by SmolVLA: 2,500 episodes, 50 per task; 5,854 recent downloads.294

Changes at test
New goal positions in ML1, or 5 held-out tasks in ML10 and ML45201
More

Depends on the protocol. MT1, MT10 and MT50 test on the goal positions seen in training. ML1 tests new goal positions; ML10 and ML45 test 5 held-out tasks.

The docs say multi-task agents are evaluated on 'the same ones seen during training'.

Scoring and access

Scored by
Success rate201
More

Rewards are for training; evaluation uses a per-task success flag, usually an object within a distance threshold (for example 5 cm) of its goal.

Score
Success rate. The protocol varies between papers.2024+1
More

An episode counts as a success if the success flag is raised at any point. Multi-task RL: one episode per training goal (50 goals) per task, averaged. Meta-RL: 3 episodes per test goal per test task after adaptation. Imitation-learning (VLA) papers: 10 episodes for each of the 50 tasks, then the plain mean of four difficulty-tier averages.

The tier grouping (easy 28, medium 11, hard 6, very hard 5) comes from Seo et al. (2023), as cited by SmolVLA and TinyVLA.

Trials
50 per task in RL papers, 10 per task in VLA papers20274+2
More

Official evaluation utility · Multi-task: one episode per each of 50 goal positions per task, 500-step horizon. Meta: 40 test goals per test task, 3 episodes each after adaptation.20

Meta-World+ experiments · 10 seeds, interquartile mean with 95% confidence intervals.2

LeRobot · Recommends 10 episodes per task (500 for MT50); default evaluation on the 'medium' group.27

FabriVLA · 10 episodes per task, 400-step horizon.16

Who runs it
Each team tests its own model Inferred330
More

No submission process or organiser evaluation.

Error bars
Sometimes reported Inferred1221+2
More

Meta-World+ reports IQM with 95% CIs over 10 seeds and MOORE reports standard deviations. The original paper's Table 1 gives means over 10 seeds without spread. SmolVLA, Evo-1 and FabriVLA give single numbers.

Leaderboard
None. Scores are only in papers. Inferred3302
More

No leaderboard on the docs site, README or Meta-World+ paper. Meta-World+ advises users to run their own baselines rather than copy numbers.

Code licence
MIT910
More

LICENSE: MIT, Copyright (c) 2019 Meta-World Team. PyPI metadata: MIT License.

Data licence
not-applicable Inferred329
More

No dataset distributed by the maintainers. The third-party LeRobot copy lerobot/metaworld_mt50 carries an Apache-2.0 tag.

Asset licence
MIT Inferred289
More

The repository's only licence file is the root MIT LICENSE; the robot and object model files sit in the same repository. Their original sources are not stated. Not legal advice.

Access
Open. It installs with pip.310
More

pip install metaworld; no registration.

Commercial use
Allowed Inferred928
More

MIT licence for code and bundled assets; no dataset. Not legal advice.

Published at
CoRL 2019, and NeurIPS 2025 Datasets and Benchmarks for Meta-World+24313

Sources 31

  1. 1Meta-World paper, full text v2 (updated version of the CoRL 2019 paper)Paper · Jun 2021 · checked 10 Oct 2026
  2. 2Meta-World+ full text v2Paper · Nov 2025 · checked 10 Oct 2026
  3. 3Meta-World READMERepository · 2026 · checked 11 Oct 2026
  4. 4SmolVLA: A vision-language-action model for affordable and efficient robotics (Table 2)Paper · Jun 2025 · checked 11 Oct 2026
  5. 5A Generalist Agent (Gato)Paper · May 2022 · checked 11 Oct 2026
  6. 6Rho: A Foundation for Efficiently Adaptable VLA ModelsPaper · Sep 2026 · checked 11 Oct 2026
  7. 7Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic AlignmentPaper · Nov 2025 · checked 11 Oct 2026
  8. 8What Are We Actually Benchmarking in Robot Manipulation? (2026 audit)Paper · Jun 2026 · checked 10 Oct 2026
  9. 9Meta-World LICENSE (MIT, Copyright (c) 2019 Meta-World Team)Repository · 2019 · checked 11 Oct 2026
  10. 10PyPI: metaworld (versions, licence, mujoco==3.3.0 requirement)Repository · 28 Jun 2026 · checked 11 Oct 2026
  11. 11TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic ManipulationPaper · Sep 2024 · checked 11 Oct 2026
  12. 12πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (v3, Table 6)Paper · Oct 2025 · checked 10 Oct 2026
  13. 13AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile ManipulationPaper · Mar 2026 · checked 11 Oct 2026
  14. 14Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action ModelPaper · May 2026 · checked 11 Oct 2026
  15. 15LA4VLA: Learning to Act without Seeing via Language-Action PretrainingPaper · Jun 2026 · checked 11 Oct 2026
  16. 16FabriVLA: A Lightweight Vision-Language-Action Model with Conformal Action Chunk Uncertainty (v3)Paper · Jul 2026 · checked 11 Oct 2026
  17. 17PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
  18. 18X2Real Technical Report (related work on simulation benchmarks)Paper · Sep 2026 · checked 10 Oct 2026
  19. 19Meta-World docs: Benchmark Descriptions (benchmark_descriptions.md)Repository · 2026 · checked 11 Oct 2026
  20. 20Meta-World docs: Evaluation (evaluation.md)Repository · 2026 · checked 11 Oct 2026
  21. 21Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts (MOORE)Paper · Nov 2023 · checked 11 Oct 2026
  22. 22One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy (OneWM)Paper · May 2026 · checked 11 Oct 2026
  23. 23Meta-World releases (v2.0.0, 3.0.0, v3.1.1) with release notesRepository · 28 Jun 2026 · checked 11 Oct 2026
  24. 24Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning (arXiv abstract page; v1 2019-10-24, v2 2021-06-14)Paper · Oct 2019 · checked 10 Oct 2026
  25. 25GitHub API: Farama-Foundation/Metaworld (stars, forks, open issues, pushed)Index · 11 Oct 2026 · checked 11 Oct 2026
  26. 26Meta-World commit historyRepository · 9 Oct 2026 · checked 11 Oct 2026
  27. 27LeRobot documentation: Meta-World (metaworld.mdx) and its commit historyRepository · Jun 2026 · checked 11 Oct 2026
  28. 28Meta-World repository tree (assets folders; licence files)Repository · 2026 · checked 11 Oct 2026
  29. 29Hugging Face Hub API record for lerobot/metaworld_mt50 (licence tag, downloads)Index · 11 Oct 2026 · checked 11 Oct 2026
  30. 30Meta-World documentation site (HTTP Last-Modified 2026-09-12)Official site · Sep 2026 · checked 11 Oct 2026
  31. 31Meta-World+: An Improved, Standardized, RL Benchmark (arXiv abstract page; v1 2025-05-16, v2 2025-11-21)Paper · May 2025 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: Meta-World paper v2 and Meta-World+ (no real-robot experiments); README and docs; PolaRiS (cites Meta-World, no measurement); X2Real (related work only); SmolVLA and FabriVLA (real-robot tests on separate tasks). A dedicated web search could not run on 2026-10-11 because the shared search budget was used up; the earlier inventory search (2026-10-10) also found no paired study.

leaderboard: README, docs site, Meta-World+ paper.

top_score: One web search for 2026 tier-average results (2026-10-10), then the papers listed in top_score; FabriVLA's comparison table.

license_assets: Repository tree (only the root LICENSE), asset XML headers.

Change history

  1. Created at full depth from the checked basic entry and primary sources. sim_to_real changed from 'none-found' to 'demonstrated'; added the VLA tier-average score series, RL protocol results and version history. Most Meta-World sources were opened on 2026-10-11 (local time) and carry that date.
  2. Published as a full entry.