Meta-World
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Meta-World is a set of 50 simulated tabletop tasks for one Sawyer robot arm. It is used to test multi-task and meta reinforcement learning (learning to adapt quickly to new tasks), and recent vision-language-action (VLA) papers also report results on it.123
What a score here does not tell you Inferred
- How well a policy (the robot's control model) will do on a real robot.We found no study that compares Meta-World scores with real-robot scores for the same policies.
- How well a policy copes with visual changes.All tasks use one fixed scene. Policies are usually tested on goal positions seen in training.
- Whether results from different papers can be compared.Papers use different reward versions and different rules for averaging scores.
Comparisons with real robots
Tried on real robots, not compared Inferred124+3
Details
Neither the original paper nor Meta-World+ has real-robot experiments. PolaRiS cites Meta-World among simulation benchmarks that fail to capture real-world visual complexity, without measuring it. The basic entry graded this 'none-found'; we use 'demonstrated' because the taxonomy defines it as 'some policies also ran on real robots'.
Our assessment Opinion
Compare numbers only when they use the same protocol and version.
Reasoning
A 'Meta-World score' can come from multi-task reinforcement learning (RL) with V1 or V2 rewards, from meta-RL on held-out tasks, or from an imitation-learning tier average. These are different measurements. Compare numbers only when they use the same protocol, version and averaging rule.
Confidence: high
A high score shows that a policy can do these tasks in one fixed simulated scene.
Reasoning
A high score shows that a policy can handle 50 short tabletop tasks in one fixed simulated scene, usually on goal positions seen in training. It says little about performance on real robots or about how well the policy copes with visual changes.
Confidence: medium
Meta-World is actively maintained. Papers should state the versions they used so that others can reproduce their results.
Reasoning
Meta-World is actively maintained and has numbered versions. This makes it a reasonable choice for RL research that needs reproducible results, as long as papers state the package version, the reward version and the MuJoCo version.
Confidence: medium
Read the score for each tier as well as the average.
Reasoning
For VLA results, read the score for each difficulty tier. The headline average gives the 5 very hard tasks as much weight as the 28 easy ones. Small gains in the average can come from a few tasks.
Confidence: medium
Known problems 5
Protocol details differ between the paper, the docs and later papers
Sources disagree on whether goal positions are fixed or random. They also give different numbers of test goals.11920+2
Details
The original paper and the benchmark docs say MT10 and MT50 goal positions are fixed (the docs carry a 'TODO: check this'), while the evaluation docs cycle through 50 goal positions per task; MOORE and PaCo report 'MT10-rand' with random goals. ML1 is described with 50 held-out positions (paper), 10 (benchmark docs) and 40 (evaluation docs). The original paper reports the average of the maximum success reached during training. Horizons differ (500 steps in the docs, 400 in FabriVLA).
VLA papers average four difficulty tiers instead of 50 tasks
VLA papers take the plain mean of four difficulty-tier averages. This gives the 5 very hard tasks as much weight as the 28 easy ones. Inferred41612+2
Details
Recent VLA papers report the plain mean of four difficulty-tier averages (28 easy, 11 medium, 6 hard, 5 very hard tasks), so each very hard task weighs about 5.6 times as much as an easy one. For SmolVLA (0.45B) the tier mean is 57.3%, while a task-weighted mean of the same tiers is about 66.8% (our arithmetic). The same model also appears with different numbers: π0 at 47.9, 50.5, 50.8 and 47.91; SmolVLA at 57.3 or 68.2 depending on model size.
Scores on some protocols are approaching 100%
In VLA papers, the average over the four tiers reaches 90%. Meta-learning scores stay near 30% to 40%.111615+2
Details
On the VLA tier-average protocol, reported scores rose from 31.6% (TinyVLA, 2024-09) to 90.0% (FabriVLA, 2026-07), with individual tiers at up to 100%. MT10 reaches 88.7% (MOORE, 2024). Meta-learning remains far from the ceiling: Meta-World+ re-runs with V2 rewards give 25.7% to 36.76% on ML10 and ML45 (MAML and RL2).
Results depend on the MuJoCo version
Newer versions of MuJoCo (the physics simulator) change how contacts between objects are represented. For this reason, v3.1.1 requires MuJoCo 3.3.0.2310
Details
v3.1.1 pins mujoco==3.3.0 because newer MuJoCo versions change how contacts between objects are represented. The same release fixed the button-press dense rewards. Results under other MuJoCo versions, or before the fix, may differ.
Two reward versions made published results incomparable
The original V1 rewards (the feedback signal used in reinforcement learning) were replaced by V2 rewards without documentation. The two versions give different scores, so older results cannot be compared.221
Details
Meta-World+ reports that the original 'V1' dense rewards were overwritten by 'V2' rewards without documentation, so papers used different rewards. In its re-runs every multi-task method scored higher with V2: PaCo reached 26.2% on MT10 with V1 and 73.6% with V2. Published numbers disagree too: PaCo's MT10 is 71.6 in its own paper and 85.4 as reported by MOORE; PCGrad's published 90% on MT10 could not be reproduced (70.5%), likely because it used a Meta-World version no longer available.
Details
About
- What it is
- Benchmark12
More
A suite of 50 tasks with defined multi-task and meta-learning evaluation protocols.
- Built by
- Stanford University, UC Berkeley, Columbia University, University of Southern California, Robotics at Google, Farama Foundation (maintainer since 2023)1232
More
Original authors (2019) · Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Avnish Narayan, Hayden Shively, Adithya Bellathur, Karol Hausman, Chelsea Finn, Sergey Levine: Stanford, UC Berkeley, Columbia, USC, Robotics at Google.1
Farama Foundation · Maintains the code since 2023; release v2.0.0 (2023-06-16) is 'the code before the Farama Foundation started maintaining Meta-World'.23
Meta-World+ authors (2025) · Toronto Metropolitan University, University of Surrey, Hamburg University of Technology, Columbia, Google DeepMind, USC, Farama Foundation, Université de Montréal / Mila.2
- Released
- October 2019, at CoRL 20192425
More
arXiv v1 on 2019-10-24; CoRL 2019. The paper was updated as arXiv v2 on 2021-06-14.
GitHub repository created 2019-09-09.
- Version
- 3.1.1, released June 20262310
More
Environment names now end in '-v3'.
v2.0.0 (2023-06-16) · Code as it was before Farama took over maintenance.23
3.0.0 (2025-06-14) · MuJoCo Python bindings replace mujoco-py; Gymnasium 1.0 API; the original 'V1' reward functions exposed again as an option; V1 environments removed.23
v3.1.1 (2026-06-28) · Pins mujoco==3.3.0 because newer MuJoCo changes how contacts are represented; fixes button-press dense rewards; adds a camera. Release notes say v3.1.0 was skipped, but PyPI lists 3.1.0 (uploaded 2026-06-26).2310
Meta-World+ (paper, 2025) · Describes the v3 changes, adds MT25 and ML25 task sets and custom task sets.2
Setup
- Runs in
- Simulation1
- Robot
- One arm1
- Robot model
- Sawyer (simulated)12
More
Actions: 3D end-effector displacement plus a gripper value. Observations: 39-dimensional state including object and goal positions; goal zeroed for meta-learning.
- Tasks
- 50 tasks128
More
50 tasks; task sets MT1, MT10, MT50, ML1, ML10, ML45, plus MT25 and ML25 since 2025
The repository has 50 '_v3' task environment files (counted by us).
- Changes at test
- New goal positions in ML1, or 5 held-out tasks in ML10 and ML45201
More
Depends on the protocol. MT1, MT10 and MT50 test on the goal positions seen in training. ML1 tests new goal positions; ML10 and ML45 test 5 held-out tasks.
The docs say multi-task agents are evaluated on 'the same ones seen during training'.
Scoring and access
- Scored by
- Success rate201
More
Rewards are for training; evaluation uses a per-task success flag, usually an object within a distance threshold (for example 5 cm) of its goal.
- Score
- Success rate. The protocol varies between papers.2024+1
More
An episode counts as a success if the success flag is raised at any point. Multi-task RL: one episode per training goal (50 goals) per task, averaged. Meta-RL: 3 episodes per test goal per test task after adaptation. Imitation-learning (VLA) papers: 10 episodes for each of the 50 tasks, then the plain mean of four difficulty-tier averages.
The tier grouping (easy 28, medium 11, hard 6, very hard 5) comes from Seo et al. (2023), as cited by SmolVLA and TinyVLA.
- Trials
- 50 per task in RL papers, 10 per task in VLA papers20274+2
More
Official evaluation utility · Multi-task: one episode per each of 50 goal positions per task, 500-step horizon. Meta: 40 test goals per test task, 3 episodes each after adaptation.20
Meta-World+ experiments · 10 seeds, interquartile mean with 95% confidence intervals.2
LeRobot · Recommends 10 episodes per task (500 for MT50); default evaluation on the 'medium' group.27
FabriVLA · 10 episodes per task, 400-step horizon.16
- Who runs it
- Each team tests its own model Inferred330
More
No submission process or organiser evaluation.
- Error bars
- Sometimes reported Inferred1221+2
More
Meta-World+ reports IQM with 95% CIs over 10 seeds and MOORE reports standard deviations. The original paper's Table 1 gives means over 10 seeds without spread. SmolVLA, Evo-1 and FabriVLA give single numbers.
- Leaderboard
- None. Scores are only in papers. Inferred3302
More
No leaderboard on the docs site, README or Meta-World+ paper. Meta-World+ advises users to run their own baselines rather than copy numbers.
- Data licence
- not-applicable Inferred329
More
No dataset distributed by the maintainers. The third-party LeRobot copy lerobot/metaworld_mt50 carries an Apache-2.0 tag.
- Asset licence
- MIT Inferred289
More
The repository's only licence file is the root MIT LICENSE; the robot and object model files sit in the same repository. Their original sources are not stated. Not legal advice.
Sources 31
- 1Meta-World paper, full text v2 (updated version of the CoRL 2019 paper)Paper · Jun 2021 · checked 10 Oct 2026
- 2Meta-World+ full text v2Paper · Nov 2025 · checked 10 Oct 2026
- 3Meta-World READMERepository · 2026 · checked 11 Oct 2026
- 4SmolVLA: A vision-language-action model for affordable and efficient robotics (Table 2)Paper · Jun 2025 · checked 11 Oct 2026
- 5A Generalist Agent (Gato)Paper · May 2022 · checked 11 Oct 2026
- 6Rho: A Foundation for Efficiently Adaptable VLA ModelsPaper · Sep 2026 · checked 11 Oct 2026
- 7Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic AlignmentPaper · Nov 2025 · checked 11 Oct 2026
- 8What Are We Actually Benchmarking in Robot Manipulation? (2026 audit)Paper · Jun 2026 · checked 10 Oct 2026
- 9Meta-World LICENSE (MIT, Copyright (c) 2019 Meta-World Team)Repository · 2019 · checked 11 Oct 2026
- 10PyPI: metaworld (versions, licence, mujoco==3.3.0 requirement)Repository · 28 Jun 2026 · checked 11 Oct 2026
- 11TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic ManipulationPaper · Sep 2024 · checked 11 Oct 2026
- 12πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (v3, Table 6)Paper · Oct 2025 · checked 10 Oct 2026
- 13AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile ManipulationPaper · Mar 2026 · checked 11 Oct 2026
- 14Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action ModelPaper · May 2026 · checked 11 Oct 2026
- 15LA4VLA: Learning to Act without Seeing via Language-Action PretrainingPaper · Jun 2026 · checked 11 Oct 2026
- 16FabriVLA: A Lightweight Vision-Language-Action Model with Conformal Action Chunk Uncertainty (v3)Paper · Jul 2026 · checked 11 Oct 2026
- 17PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
- 18X2Real Technical Report (related work on simulation benchmarks)Paper · Sep 2026 · checked 10 Oct 2026
- 19Meta-World docs: Benchmark Descriptions (benchmark_descriptions.md)Repository · 2026 · checked 11 Oct 2026
- 20Meta-World docs: Evaluation (evaluation.md)Repository · 2026 · checked 11 Oct 2026
- 21Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts (MOORE)Paper · Nov 2023 · checked 11 Oct 2026
- 22One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy (OneWM)Paper · May 2026 · checked 11 Oct 2026
- 23Meta-World releases (v2.0.0, 3.0.0, v3.1.1) with release notesRepository · 28 Jun 2026 · checked 11 Oct 2026
- 24Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning (arXiv abstract page; v1 2019-10-24, v2 2021-06-14)Paper · Oct 2019 · checked 10 Oct 2026
- 25GitHub API: Farama-Foundation/Metaworld (stars, forks, open issues, pushed)Index · 11 Oct 2026 · checked 11 Oct 2026
- 26Meta-World commit historyRepository · 9 Oct 2026 · checked 11 Oct 2026
- 27LeRobot documentation: Meta-World (metaworld.mdx) and its commit historyRepository · Jun 2026 · checked 11 Oct 2026
- 28Meta-World repository tree (assets folders; licence files)Repository · 2026 · checked 11 Oct 2026
- 29Hugging Face Hub API record for lerobot/metaworld_mt50 (licence tag, downloads)Index · 11 Oct 2026 · checked 11 Oct 2026
- 30Meta-World documentation site (HTTP Last-Modified 2026-09-12)Official site · Sep 2026 · checked 11 Oct 2026
- 31Meta-World+: An Improved, Standardized, RL Benchmark (arXiv abstract page; v1 2025-05-16, v2 2025-11-21)Paper · May 2025 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: Meta-World paper v2 and Meta-World+ (no real-robot experiments); README and docs; PolaRiS (cites Meta-World, no measurement); X2Real (related work only); SmolVLA and FabriVLA (real-robot tests on separate tasks). A dedicated web search could not run on 2026-10-11 because the shared search budget was used up; the earlier inventory search (2026-10-10) also found no paired study.
leaderboard: README, docs site, Meta-World+ paper.
top_score: One web search for 2026 tier-average results (2026-10-10), then the papers listed in top_score; FabriVLA's comparison table.
license_assets: Repository tree (only the root LICENSE), asset XML headers.
Change history
- Created at full depth from the checked basic entry and primary sources. sim_to_real changed from 'none-found' to 'demonstrated'; added the VLA tier-average score series, RL protocol results and version history. Most Meta-World sources were opened on 2026-10-11 (local time) and carry that date.
- Published as a full entry.