HumanoidBench

HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

How to read this picture

HumanoidBench is a set of 27 simulated tasks for a humanoid robot with two hands. It tests reinforcement learning methods (which learn by trial and error from a reward) and scores them by the reward they collect.123

Sources
Last checked 10 Oct 2026Full entry41 of 55 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation12
Checked against real robots
Not checked
Skill
Walking and balance
Robot
Humanoid, Robot hand1
Used by
20 reinforcement learning papers, up to October 2026456+7
170 citations
Licence
MIT14
Commercial use: allowed

What a score here does not tell you Inferred

  1. How well a controller will do on a real robot.HumanoidBench runs only in simulation. Its main robot model cannot be built as a real robot.
  2. Whether the motion is natural and safe.A high reward can come from odd postures.
  3. How results from different papers compare.Papers differ in task sets, in whether the robot has hands and in the number of training steps.

Comparisons with real robots

Not checked Inferred1157

Details

The paper lists sim-to-real transfer as future work. A user asked about sim-to-real support on 2026-03-01 (issue #68) without reply. FastTD3 shows a real Booster T1 robot, but its policy was trained in MuJoCo Playground, not HumanoidBench. See searched.

Our assessment Opinion

It is a simulation test for learning methods. It does not test real robots.

Reasoning

HumanoidBench tests reinforcement-learning methods in simulation: how fast and how far a method raises the reward. It says little about real humanoid robots, because the robot model is not buildable and nothing has been checked on hardware.

Confidence: high

Look at the results for each task instead of averages.

Reasoning

Read per-task results. Averages hide that walking-type tasks are close to the threshold while stairs, hurdles, the maze and most manipulation tasks are far from it.

Confidence: medium

Check the settings before comparing papers.

Reasoning

Before comparing two papers, check whether hands were on, which task subset was used, the step budget, the action repeat and the normalisation.

Confidence: high

High reward does not guarantee good motion.

Reasoning

A high return is not proof of good motion. Reward-maximising behaviour can be jerky or unsafe, so videos or extra measures of smoothness are worth asking for.

Confidence: medium

Known problems 6

  1. Bug reports remain open with no reply from the maintainer

    Bug reports have had no replies, including one about heavy packages in the truck task. Inferred161718+3

    Details

    Several 2025 reports are unanswered. Issue #40 (2025-04) could not reproduce the authors' window-task results. Issue #60 (2025-10) reports that truck-task packages have no density set. Our reading of truck.xml confirms this, so MuJoCo's default of 1000 kg per cubic metre applies, which makes the smallest package about 24 kg; the separate package task sets a density of 5. The maintainer's last issue comment was on 2024-09-24.

  2. Papers test different versions of the benchmark

    Papers differ in whether the robot has hands, which tasks they use, how many training steps they allow and how they normalise scores.568+4

    Details

    The main benchmark uses the H1 with hands (h1hand-*). Many algorithm papers instead use 14 locomotion tasks without hands (SimBa, SimbaV2, FlashSAC), sometimes with an action repeat of 2. BRC and EZ-M use a 9-task no-hand set and a 14-task set with hands. FastTD3 uses 39 tasks with many parallel simulations. Step budgets range from 1M steps to hours of parallel training, and three different normalisations are in use. Scores from different papers are therefore not comparable without checking these settings.

  3. Success thresholds are rough

    The success lines are rough levels of total reward. They do not check that the task was done. Inferred1229+1

    Details

    The paper says its dashed success lines 'qualitatively indicate task success'. They are fixed return levels in code, inherited by related tasks (run, crawl, stair, slide and hurdle all use walk's 700). Passing the line does not check that the task was completed in a defined way, and some later papers print different thresholds (EZ-M lists 800 for h1hand-crawl). The repository's published PPO curves also look like fillers: the paper says PPO ran on 4 tasks, and the JSON's 31 PPO entries are copies of 4 distinct curve sets.

  4. Simple locomotion tasks are close to solved

    Most basic walking tasks are now passed. The harder tasks are not. Inferred2396

    Details

    In 2024 the best original baseline passed the threshold on 3 of 31 environments. By 2026, within 1M training steps, at least one published method passes on 6 of 9 no-hand locomotion tasks and 7 of 14 with-hand tasks (EZ-M tables). Stairs, hurdles, the maze, balance boards and reach stay well below, and most manipulation tasks are rarely reported. The suite is not saturated as a whole.

  5. The simulated robot does not exist

    The main robot model cannot be built, and nothing was run on hardware.115

    Details

    The main robot is a Unitree H1 with two Shadow Hands whose forearms were removed; the paper calls this 'not currently a realistic model'. No HumanoidBench policy has been tested on hardware, and a 2026 request for sim-to-real support has no reply.

  6. Maximising the reward can give unrealistic motion

    A high reward can come from postures a real robot should avoid.1317

    Other view: The critics say episodes last 2 seconds. The code runs them for up to 20 seconds.2425

    Details

    Meser and others (TU Darmstadt and DFKI, 2024-08) ran model predictive control on stand, walk and push. They found that optimising HumanoidBench's rewards gives undesirable and unrealistic behaviour: in push, the reward drives the robot into unrecoverable postures to reach the box fast. Adding posture and smoothness terms gave higher HumanoidBench scores with steadier motion. The original paper also reports failures such as clinging to the high bar and colliding with hurdles instead of jumping. FastTD3 (2025-05) found in another simulator that one reward function can give a natural gait with one algorithm and an undeployable gait with another.

    Part of the critique does not match the code. Meser and others say walk and stand episodes last 2 seconds; the code runs up to 1000 control steps of 0.02 s each, which is 20 seconds.

Details

About

What it is
Benchmark1
More

The paper presents a fixed task suite with reward functions, success thresholds and baseline results.

Built by
UC Berkeley, Yonsei University12
More

Authors: Carmelo Sferrazza, Dun-Ming Huang, Xingyu Lin, Youngwoon Lee, Pieter Abbeel, all at UC Berkeley; Youngwoon Lee also lists Yonsei University. The VLGE claim of a KAIST affiliation is not supported by the paper.

UC Berkeley · All five authors (Robot Learning Lab, per the LICENSE).114

Yonsei University · Second affiliation of Youngwoon Lee.12

Released
March 2024, at RSS 2024262728
More

arXiv v1 on 2024-03-15; code repository created 2024-03-18. Published at RSS 2024 (Delft, July 2024).

RSS DOI 10.15607/RSS.2024.XX.061. arXiv v2 is dated 2024-06-18.

Version
0.2.0, from July 20242930
More

Git tags v0.1.0 and v0.2.0, both published as releases on 2024-07-02. v0.2.0 added the Unitree G1 robot. setup.py says version 0.2.

Not on PyPI as far as we checked; installed from source with pip install -e.

Last update
September 2025. Only a licence note was added.31
More

2025-09-18: note added to the LICENSE file. Last code change 2025-05-20 (lighter dependencies).

The 2025-09 commits only add a BAIR Commons note to the LICENSE. Code merges in April and May 2025 came from FastTD3 authors (rendering switch, print removal).

Status
No code changes since May 2025. Still in active use. Inferred3121
More

No code change since May 2025. Widely used as a reinforcement-learning test suite.

Last code change 2025-05-20; last commit of any kind 2025-09-18, over a year before the check date. The maintainer's last issue comment is dated 2024-09-24; 26 issues and pull requests are open, including bug reports from 2025. The basic entry said 'maintained'; changed by the one-year rule.

Setup

Runs in
Simulation12
Simulator
MuJoCo 3.1.613024+1
More

MuJoCo (pinned 3.1.6), 0.002 s physics step, control at 50 Hz. MuJoCo MJX used to pre-train low-level reaching skills.

Each control step runs 10 physics steps; episodes last up to 1000 control steps (20 s of simulated time). The paper reports over 1,000 frames per second on one CPU for the default model.

Robot
Humanoid, Robot hand1
Robot model
Unitree H1 with two Shadow Hands13
More

Main robot: Unitree H1 with two Shadow Hands, 61 actuated joints and 75 degrees of freedom. Also provided: Unitree G1 with three-finger hands, Agility Digit, Robotiq 2F-85 gripper, Unitree H1 hands.

The authors removed the Shadow Hands' forearms to make the robot more human-shaped and say 'this is not currently a realistic model'. H1 runs with position control; G1 and Digit are registered with torque control.

Setting
Mixed Inferred13
More

Read from the task list.

Tasks
27 tasks and 31 environments13
More

27 tasks: 12 locomotion and 15 manipulation. 31 environments when easy and hard variants count separately.

The README's main list has 31 IDs (h1hand-* plus h1strong-highbar_hard-v0). The same tasks also exist without hands (h1-*, 20 IDs) and for the G1 robot (30 IDs).

Training data
None. Agents learn from the reward.13
More

None. Agents learn from the reward by trial and error.

The repository ships pre-trained low-level reaching policies for the hierarchical baseline, not demonstrations.

Changes at test
Only targets and start positions are randomised. Inferred1
More

Train and test use the same task. Targets and start positions are randomised in some tasks.

For example, push and reach use random targets, sit_hard randomises the robot's pose, cube uses random target orientations and basketball throws from random directions. There is no held-out test set; agents learn and are scored in the same environment.

Scoring and access

Scored by
Reward1
More

Episode return (summed reward), shown as learning curves against environment steps, with a per-task success threshold.

Score
Reward compared with a fixed threshold for each task122
More

Each task has a hand-written reward. For walking-type tasks each step's reward is at most 1, so an episode of 1000 steps scores at most 1000. Papers plot this return against training steps. A dashed line marks a fixed per-task threshold, for example 700 for walk, 800 for stand, 1200 for maze, 12000 for reach and 4 for kitchen. The paper says the lines 'qualitatively indicate task success'.

Thresholds are the success_bar values in the environment code. Later papers normalise in different ways: return divided by threshold (SimBa, SimbaV2, FlashSAC); (return minus random score) divided by (threshold minus random score) (BRC, EZ-M); or return relative to one baseline (EfficientTDMPC).

Trials
Usually 3 seeds. The number of training steps varies.169+1
More

Original paper: 3 seeds, about 48 hours of training per run (about 2M steps for TD-MPC2, 10M for DreamerV3, SAC and PPO). Common later protocol: 1M environment steps.

SimBa / SimbaV2 · 14 locomotion tasks without hands (h1-*), 1M environment steps with action repeat 2, 95% intervals; 10 seeds for SAC-based runs in SimBa.56

BRC and EZ-M · HumanoidBench-Medium (9 no-hand locomotion tasks) and HumanoidBench-Hard (14 tasks with hands), 1M steps, no action repeat, 3 seeds.89

FastTD3 · 39 tasks with many parallel simulations, 3 runs; notes that SimbaV2's action repeat of 2 'is unusual for joint position control'.7

Who runs it
Each team tests its own model Inferred23
More

No organiser runs submissions. The repository publishes the authors' own training curves as JSON so others can compare without re-running.

Error bars
Usually reported Inferred169+1
More

The original paper shades one standard deviation over 3 seeds; SimbaV2 and EZ-M report 95% intervals; FastTD3 shades standard deviation. With 3 seeds these intervals are wide.

Leaderboard
None. Scores are only in papers.23
More

No results table on the project page or in the README; only the authors' baseline curves in logs/main_results.json.

Code licence
MIT14
More

LICENSE: MIT, copyright 2024 Robot Learning Lab at UC Berkeley. It also reproduces licences of bundled code: jaxrl_m, DreamerV3 and TD-MPC2 (MIT), purejaxrl (Apache-2.0), MuJoCo (Apache-2.0). GitHub reports 'NOASSERTION' because the file holds several licences.

Data licence
not applicable Inferred3
More

No dataset is shipped. Pre-trained reaching-policy weights sit in the repository under its LICENSE.

Asset licence
Robot models: Unitree (BSD-3-Clause), Shadow Hand (Apache-2.0), Digit (MIT), Robotiq 2F-85 (BSD-style, ROS-Industrial); robosuite textures (MIT).14
More

All in the LICENSE file since the first commit (2024-03-18). This corrects the inventory note that asset licences were not restated.

Access
Open. The code is on GitHub.3232
More

Code on GitHub, installed from source. The project page humanoid-bench.github.io works; the README's website link (sferrazza.cc/humanoidbench_site/) returns HTTP 404.

Commercial use
Allowed Inferred14
More

Code (MIT) and bundled robot models (BSD-3-Clause, Apache-2.0, MIT, BSD-style) are permissive with attribution. Not legal advice.

Sources 32

  1. 1HumanoidBench paper, full text (arXiv v2)Paper · Jun 2024 · checked 10 Oct 2026
  2. 2HumanoidBench project pageOfficial site · 2024 · checked 10 Oct 2026
  3. 3HumanoidBench GitHub READMERepository · Sep 2024 · checked 10 Oct 2026
  4. 4Semantic Scholar API record for arXiv:2403.10506Index · 10 Oct 2026 · checked 10 Oct 2026
  5. 5SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning (Appendix H.3)Paper · Oct 2024 · checked 10 Oct 2026
  6. 6Hyperspherical Normalization for Scalable Deep Reinforcement Learning (SimbaV2; Table 1)Paper · Feb 2025 · checked 10 Oct 2026
  7. 7FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid ControlPaper · May 2025 · checked 10 Oct 2026
  8. 8Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners (BRC)Paper · May 2025 · checked 10 Oct 2026
  9. 9Scaling Tasks, Not Samples: Mastering Humanoid Control through Multi-Task Model-Based RL (EZ-M; Appendix B)Paper · Mar 2026 · checked 10 Oct 2026
  10. 10FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot ControlPaper · Apr 2026 · checked 10 Oct 2026
  11. 11FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid ControlPaper · Mar 2026 · checked 10 Oct 2026
  12. 12EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous ControlPaper · May 2026 · checked 10 Oct 2026
  13. 13MuJoCo MPC for Humanoid Control: Evaluation on HumanoidBenchPaper · Aug 2024 · checked 10 Oct 2026
  14. 14HumanoidBench LICENSE file (MIT plus third-party licences)Repository · 18 Sep 2025 · checked 10 Oct 2026
  15. 15Issue #68: Will HB support sim to real in the future release? (no reply)Repository · Mar 2026 · checked 10 Oct 2026
  16. 16Issue #40: Unable to reproduce h1hand_window results (no reply)Repository · Apr 2025 · checked 10 Oct 2026
  17. 17Issue #60: No density or mass on liftable packages in truck environment (no reply)Repository · Oct 2025 · checked 10 Oct 2026
  18. 18truck.xml task asset (package geoms without density)Repository · 2024 · checked 10 Oct 2026
  19. 19package.xml task asset (density 5)Repository · 2024 · checked 10 Oct 2026
  20. 20MuJoCo XML reference: geom density default 1000Official site · 2026 · checked 10 Oct 2026
  21. 21HumanoidBench GitHub issues and pull requestsRepository · Jun 2026 · checked 10 Oct 2026
  22. 22HumanoidBench environment code (success_bar values per task)Repository · 2024 · checked 10 Oct 2026
  23. 23HumanoidBench logs/main_results.json (authors' training curves for 31 environments)Repository · Apr 2024 · checked 10 Oct 2026
  24. 24HumanoidBench tasks.py (frame_skip 10, max_episode_steps 1000)Repository · 2024 · checked 10 Oct 2026
  25. 25h1hand_pos_walk.xml (physics timestep 0.002 s)Repository · 2024 · checked 10 Oct 2026
  26. 26HumanoidBench arXiv abstract page (v1 2024-03-15, v2 2024-06-18)Paper · Mar 2024 · checked 10 Oct 2026
  27. 27GitHub API: carlosferrazza/humanoid-bench (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
  28. 28HumanoidBench, Robotics: Science and Systems XX (RSS 2024) proceedings page, paper 61Paper · Jul 2024 · checked 10 Oct 2026
  29. 29HumanoidBench tags and releases (v0.1.0, v0.2.0)Repository · 2 Jul 2024 · checked 10 Oct 2026
  30. 30HumanoidBench setup.py (version 0.2, MuJoCo 3.1.6 pin)Repository · Sep 2024 · checked 10 Oct 2026
  31. 31HumanoidBench commit historyRepository · 18 Sep 2025 · checked 10 Oct 2026
  32. 32README website link (returns HTTP 404)Official site · 10 Oct 2026 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: Paper (arXiv v2, full text), project page, README and all GitHub issues; MuJoCo MPC evaluation (2408.00342); FastTD3 (2505.22642, real robot trained in MuJoCo Playground); FlashSAC (2604.04539); 31 of the 170 Semantic Scholar citing papers opened. No study compares HumanoidBench scores with real-robot results. Validity list left empty.

top_score: Authors' logs/main_results.json; SimbaV2 Table 1; EZ-M Appendix B Tables 1 and 2; FastTD3, BRC, FastDSAC, FlashSAC, WarpSAC and EfficientTDMPC texts. Papers use different subsets and normalisations, so no chart series was made.

status: Commit history, tags and releases, issue comments by the repository owner (last 2024-09-24), open pull requests.

license_assets: LICENSE file (9 sections), its history (third-party notices present since the first commit), README references.

used_by (MuJoCo MPC integration): google-deepmind/mujoco_mpc pull request #328 (closed, not merged) and the repository tree (no HumanoidBench tasks).

Change history

  1. Created at full depth from primary sources, starting from the basic entry and research/raw/inventory/academic-c.json. Changes from the basic entry: status set to dormant (no update for over a year); commercial_use set to allowed with asset licences, which the LICENSE file does list (inventory note corrected). New findings: reward critique and its episode-length error, protocol differences across RL papers, partial saturation of locomotion tasks, duplicated PPO curves in the published logs, truck package density bug.
  2. Published as a full entry.