AgiBot World
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
AgiBot World is a training dataset recorded with about 100 real AgiBot robots. It is used to pretrain robot policies (the models that control robots), and it has no public test of its own.123
What a score here does not tell you Inferred
- How the tested policies would score on anyone else's robots.All of the tests ran on AgiBot robots at AgiBot's site.
- How a result compares with public results from other teams.There is no public test, no leaderboard and no published scoring rubric.
- How good the data is for training policies on other hardware.The gains from training on it were measured on the same robot that recorded the data.
Our assessment Opinion
AgiBot World is training data, and its tests ran only inside AgiBot. Other teams cannot run them as a benchmark.
Reasoning
AgiBot World is training data. It is not a benchmark. Its paper's scores show that pretraining on it (training on it before training for a specific task) helped AgiBot's policies on AgiBot's own tasks and robots. Nobody else can run those tests, so the scores cannot be compared with other papers.
Confidence: high
The test that compares AgiBot World with other data favours AgiBot's own robot and site.
Reasoning
The comparison with Open X-Embodiment (another large robot dataset) was run on the same robot and in the same facility that produced AgiBot World's data. A gain there may come from the quality of the data or from testing on the robot and site that recorded it. An outside test on other robots would be needed to separate the two.
Confidence: medium
When you quote its size, say whether you mean episodes or sub-task trajectories.
Reasoning
When you quote its size, give the unit. It holds about 160,000 episodes, or about 1 million sub-task trajectories. The headline counts and hours differ between the paper, the README, the dataset cards and later papers.
Confidence: high
The claim about how results grow with more data is based on three data points.
Reasoning
The scaling result (Pearson r = 0.97 for a power-law fit) is based on three pretraining sizes. It shows that more data helped in these tests. It is weak evidence for a general scaling law (a fixed rule linking the amount of data to performance).
Confidence: high
Known problems 6
The camera and action fields have known problems
The maintainers have acknowledged errors in the camera calibration. They also confirmed that some action fields were copied from the robot state. The fixes so far are partial.111213+2
Details
In 2025-03 a maintainer confirmed errors in the camera extrinsic parameters and promised corrected data; the 2025-04-14 Alpha update corrected 'some data'. Users still reported misaligned projections in 2025-11 (issue #122) and 2026-02 (issue #28, open). The end-effector pose fields under 'action' equal those under 'state' in whole episodes; a maintainer said some VR controller poses were not recorded and were copied from the state, recommended joint representations, and said GO-1 is trained to predict the state as the action. A user reported frame and video misalignment in 9 tasks after conversion to LeRobot 3.0 (issue #149, open, no maintainer reply).
The evaluation cannot be re-run by others
The paper's tests ran at AgiBot's site and used a scoring rubric that is not published. No one else can repeat them. Inferred11617+1
Details
The paper's tests use 6 tasks staged in AgiBot's facility, AgiBot G1 robots, task-specific fine-tuning data and an unpublished partial-credit rubric. No simulator, leaderboard or test server exists for the dataset. The Hugging Face card still lists 'AgiBot World Colosseum: Comprehensive platform (expected release date: 2025)' as a to-do. For offline checks a maintainer suggested comparing predicted with recorded actions, adding that this may differ from real-world performance. AgiBot's later evaluation work moved to Genie Sim and the AgiBot World Challenge.
Size figures differ between sources
The trajectory and hour counts for both releases differ between the paper, the README, the dataset cards and later papers.1319+3
Details
Beta: 1,001,552 trajectories and 2,976.4 hours in the paper; 1,003,672 trajectories in the README; 'approximately one million ... episodes, totaling 2,967 hours' in AgiBot's Genie Envisioner paper. Alpha: 92,214 trajectories (README); '100,000+ trajectories ... 300 hours' (card); 236 hours (paper experiments); 474.12 hours rising to 595.31 hours after the 2025-04 update (card); 26,375 rising to 34,512 episodes (maintainer). GR00T N1 used '140,000 trajectories' of Alpha. The paper also calls Alpha 'roughly 14%' of Beta's trajectories in the text and 'around 10%' in its change log; 92,214 is 9.2% of 1,001,552.
Early data showed the faces of robot operators
Early data showed the faces of the people who operated the robots remotely. The April 2025 update anonymised the data.202119
Details
In 2025-01 a user noted that faces were not blurred in a rear camera of the sample data. AgiBot replied that it had an agreement with the teleoperators and had attempted blurring. The 2025-04-14 Alpha update says the data was anonymised to remove personal and sensitive information. Task 429 (memory-kit insertion) was removed from Alpha in 2025-01 because of a visuo-tactile sensor setup problem, to be re-collected.
The figure of 1 million trajectories counts sub-task segments
The figure of 1 million trajectories counts sub-task pieces cut from about 160,000 recorded episodes.21
Details
A maintainer explained that the dataset holds about 160,000 episodes, each split into several sub-task 'trajectories' (action slices), which gives the 1 million figure. The paper's comparison table sets this 1M+ against counts for other datasets such as DROID (76k) and RoboMIND (55k) without saying the unit differs. Counted as recorded episodes, AgiBot World is about 160,000.
Completion scores are presented as success rates
The paper's abstract and AgiBot's website describe scores that give partial credit as success rates.122
Details
Figure 5 reports normalized completion scores with partial credit (average RDT-1B 0.46, GO-1 0.78). AgiBot's GO-1 page describes the same numbers as 'increasing success rates by 32% (46% → 78%)', and the abstract says GO-1 achieves 'over 60% success rate on complex tasks'. The abstract's '30%' improvement over Open X-Embodiment is an absolute gain of 0.30 in completion score (0.47 to 0.77), not a relative gain.
Details
About
- What it is
- Dataset Inferred13
More
The paper calls AgiBot World Colosseo a 'platform' of data, models, benchmarks and ecosystem. What is released is a dataset plus a policy model (GO-1). The evaluation tasks, rubric and robots used for scoring are not released as a test others can run, so the Atlas classes it as a dataset.
- Built by
- AgiBot with HKU, SII and Shanghai AI Lab1233+1
More
Authored as 'Team AgiBot-World', authors in alphabetical order. A Chinese news report of the 2024-12-30 launch names only AgiBot (智元机器人) as the releaser.
AgiBot Inc. · Built the AgiBot G1 robot and the collection facility. Project co-lead Maoqing Yao (agibot.com address).1
The University of Hong Kong · Listed in paper v4. Not listed in v1, which named Shanghai AI Lab, AgiBot Inc. and Shanghai Innovation Institute.123
Shanghai Innovation Institute · Project co-lead Hongyang Li (sii.edu.cn address).1
Shanghai AI Lab · Project co-lead Yu Qiao (pjlab.org.cn address).1
OpenDriveLab (code host) · The code repository and project blog sit under OpenDriveLab.325
- Released
- December 2024 (Alpha) and March 2025 (full set)31626
More
Alpha released 2024-12-30; complete set (Beta) 2025-03-01; paper on arXiv 2025-03-09.
README news and the Hugging Face gate text both give 2024-12-30 for Alpha. The paper's change log says 'Jan 2025' for the Alpha release.
- Version
- Beta (complete) and Alpha (subset)31619+1
More
Two releases on Hugging Face: AgiBot World Beta, the complete set, and AgiBot World Alpha, an earlier subset. No version tags in the code repository.
AgiBot World Beta · Released 2025-03-01. README: 1,003,672 trajectories, about 43.8 TB. Hugging Face page: 48.1 TB total file size. Files last changed 2025-10-13.31628
AgiBot World Alpha · Released 2024-12-30; sample set 2025-01-03; on OpenDataLab 2025-01-20. README: 92,214 trajectories, about 8.5 TB. 36 tasks. Updated 2025-04-14 (frame-loss episodes removed, data anonymised and compressed, some camera parameters corrected). Files last changed 2025-09-29.31929+1
Paper versions · arXiv v1 2025-03-09, v2 2025-03-13, v3 2025-04-30, v4 2025-08-04. v4 is the IROS 2025 camera-ready text and adds a comparison with π0.261
AGIBOT WORLD 2026 · A separate, newer dataset on AgiBot's G2 robot, released in phases from 2026-03. Separate Atlas record (agibot-world-2026).30
Setup
- Runs in
- Real robots1
More
The paper states that all evaluations were conducted in real-world scenarios.
- Robot
- Humanoid, Two arms, Robot hand1
More
The AgiBot G1 has two arms, a waist and a wheeled base. Teleoperators can move the base, but the paper does not say which tasks use it, so mobile-manipulator is not tagged.
- Robot model
- About 100 AgiBot G1 robots1
More
AgiBot G1: two 7-DoF arms, mobile base, adjustable waist; gripper, 6-DoF dexterous hand or gripper with visuo-tactile sensors; 8 cameras; recorded at 30 Hz. More than 100 identical units.
Teleoperation by VR headset or whole-body motion capture.
- Setting
- Staged home, shop, factory, restaurant and office settings124
More
Five domains rebuilt at full scale in one 4,000 m² facility: domestic, retail, industrial, restaurant and office
The restaurant domain has no taxonomy value, so 'mixed' is added. These are staged replicas inside AgiBot's data factory, not real homes or shops. Chinese launch report (secondary): home 40%, dining 20%, industrial 20%, retail 10%, office 10%.
- Tasks
- 217 tasks1
More
217 tasks, 87 skills
The paper's evaluations use 6 of them (Restock Bag, Table Bussing, Pour Water, Restock Beverage, Fold Shorts, Wipe Table).
- Training data
- About 1 million sub-task trajectories from about 160,000 episodes12
More
1,001,552 trajectories and 2,976.4 hours (paper). These 'trajectories' are sub-task segments of about 160,000 recorded episodes.
CONFLICTS between sources; see items and issues.i1 and issues.i2.
About 160,000 episodes · A maintainer: 'The dataset contains ~160K episodes, each divided into multiple trajectories', giving about 1M trajectories. The user counted 168,869 proprioception files but 165,745 observation videos in the 2025-03-27 Beta; the maintainer said missing videos were added on 2025-04-12.2
1,003,672 trajectories · README figure for Beta (about 43.8 TB)3
Alpha: 92,214 trajectories · README (about 8.5 TB). The Alpha card says '100,000+ trajectories ... 300 hours' and that the 2025-04 update raised total duration from 474.12 to 595.31 hours. A maintainer gave the episode count as 26,375 before and 34,512 after the update. The paper's experiments describe Alpha as 236 hours.31914+1
Failure-recovery data: about 1% · Episodes where the operator recovered from an error are kept and annotated with the cause and time.1
Collection and annotation · Human teleoperation; local check for missing frames, then annotators verify each episode against the collection standard and add task and sub-step language annotations.116
- Changes at test
- New object positions, visual distractors or new wording1
More
In the paper's tests, each task is also run in 2 unseen setups: new object positions, visual distractors, or new wording.
Paper Section V-A1. The evaluation tasks themselves are drawn from the training tasks; policies are fine-tuned on task-specific demonstrations before testing.
Scoring and access
- Scored by
- Progress score122
More
Paper: 'Each episode scores 1.0 for full success, with fractional scores for partial success.' See issues.i5.
- Score
- Partial-credit score, averaged over 10 runs per setup1
More
Each rollout scores 1.0 for full success and a fraction for partial success; scores are averaged over 10 rollouts per task, setup and method. The partial-credit rubric for each task is not published.
Values below are read from Figures 5 to 7 of paper v4 (image figures).
Pretraining data comparison · RDT pretrained on Open X-Embodiment vs AgiBot World Alpha vs Beta, then fine-tuned, on 3 tasks: average 0.47 / 0.68 / 0.77 in seen setups and 0.38 / 0.56 / 0.67 in unseen setups. On Table Bussing (seen) Alpha scored 0.65 and Beta 0.60.1
GO-1 comparison · Average over 5 tasks, 30 trials each (10 seen, 20 varied): RDT-1B 0.46, π0 0.58, GO-1 without latent planner 0.66, GO-1 0.78. Per task, GO-1 ranges from 0.60 (Restock Beverage) to 1.00 (Table Bussing).122
Scaling fit · Out-of-the-box GO-1 performance on 4 seen tasks after pretraining on about 9.2k, 92k and 1M trajectories; power-law fit with Pearson r = 0.97 over these 3 points.1
Data-quality ablation · RDT fine-tuned on Wipe Table: 482 unverified plus 528 verified trajectories ('All') 0.41 vs verified only 0.59.1
- Trials
- 10 per setup, and 30 per task for GO-11
More
GO-1 tests: 10 trials in a seen setup and 20 under variations or distractors per task.
- Who runs it
- Each team tests its own model Inferred13
More
All published scores come from the builders' own tests. No outside group has run the paper's evaluation tasks.
- Error bars
- Not reported Inferred132
More
Figures 5 to 7 show single bars without error bars or confidence intervals. The team's GO-1 blog (2025-09) says real-robot test noise can exceed real improvements and describes fixing object positions and lighting to reduce it.
- Leaderboard
- None. Scores are only in papers. Inferred31625
More
No leaderboard on the README, Hugging Face cards or project page.
- Code licence
- README: 'All the data and code within this repo are under CC BY-NC-SA 4.0'. The repo has no LICENSE file.3910
More
pyproject.toml declares license = {file = 'LICENSE'}, but no such file exists; the GitHub licence endpoint returns 404.
- Data licence
- CC-BY-NC-SA-4.011628
More
Paper, dataset cards and the Hugging Face gate text ('AgiBot World COMMUNITY LICENSE AGREEMENT') all name CC BY-NC-SA 4.0.
- Access
- Free after filling in a form with contact details16283
More
Hugging Face gate with automatic approval: name, email, country, affiliation, phone, job title and research interest, plus acceptance of the licence and the AgiBot Privacy Policy. Also on OpenDataLab.
Hub API: gated = 'auto'. The gate also records IP location.
- Commercial use
- Not allowed Inferred316
More
Reading of CC BY-NC-SA 4.0, which covers both data and code. Not legal advice.
- Published at
- IEEE/RSJ IROS 2025, pages 3549-3556, DOI 10.1109/IROS60139.2025.11247088333
More
The repo description says 'IROS 2025 Best Paper Award Finalist & IEEE TRO 2026'. Crossref has no TRO record for this paper. The TRO 2026 paper is the follow-up study 'Is Diversity All You Need for Scalable Robotic Manipulation?' (vol. 42, pp. 1872-1883, DOI 10.1109/TRO.2026.3686184), which uses AgiBot World data. The best-paper-finalist claim was not checked at an IROS source.
Sources 33
- 1AgiBot World Colosseo, full text v4 (incl. Figures 5 to 7)Paper · Aug 2025 · checked 10 Oct 2026
- 2AgiBot-World issue #79: 'Dataset Incomplete? Only 160k Trajectories Found Instead of 1M' (maintainer reply)Repository · Jul 2025 · checked 10 Oct 2026
- 3OpenDriveLab/AgiBot-World READMERepository · Oct 2025 · checked 10 Oct 2026
- 4GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (pretraining data section)Paper · Mar 2025 · checked 10 Oct 2026
- 5Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationPaper · Aug 2025 · checked 10 Oct 2026
- 6EO-1: An Open Unified Embodied Foundation Model for General Robot Control (training data table)Paper · Aug 2025 · checked 10 Oct 2026
- 7Is Diversity All You Need for Scalable Robotic Manipulation?Paper · Jul 2025 · checked 10 Oct 2026
- 8villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models (pretraining data)Paper · Jul 2025 · checked 10 Oct 2026
- 9GitHub API: OpenDriveLab/AgiBot-World (stars, forks, licence)Index · 10 Oct 2026 · checked 10 Oct 2026
- 10AgiBot-World pyproject.toml (license = file LICENSE)Repository · Sep 2025 · checked 10 Oct 2026
- 11AgiBot-World issue #28: 'Error in camera extrinsic' (open)Repository · Mar 2025 · checked 10 Oct 2026
- 12AgiBot-World issue #122: inaccurate extrinsic parameters (open)Repository · Nov 2025 · checked 10 Oct 2026
- 13AgiBot-World issue #120: 'state and action are entirely the same' (maintainer reply)Repository · Oct 2025 · checked 10 Oct 2026
- 14AgiBot-World issue #43: 'Change log requirement' (Alpha update 2025-04-14)Repository · Apr 2025 · checked 10 Oct 2026
- 15AgiBot-World issue #149: frame and video misalignment in 9 tasks after LeRobot 3.0 conversion (user report, open)Repository · Mar 2026 · checked 10 Oct 2026
- 16agibot-world/AgiBotWorld-Beta dataset card and gateRepository · 13 Oct 2025 · checked 10 Oct 2026
- 17AgiBot-World issue #42: 'How to evaluate the performance without deploying on a real robot?' (maintainer reply)Repository · Mar 2025 · checked 10 Oct 2026
- 18agibot-world.com AgiBot World Challenge 2025 page text (site bundle)Official site · 2025 · checked 10 Oct 2026
- 19agibot-world/AgiBotWorld-Alpha dataset card (incl. 'Important Notice' update log)Repository · 29 Sep 2025 · checked 10 Oct 2026
- 20AgiBotWorld-Alpha discussion #23: removing personally identifiable information (AgiBot reply)Repository · Feb 2025 · checked 10 Oct 2026
- 21AgiBotWorld-Alpha discussion #18: task 429 removed (AgiBot reply)Repository · Jan 2025 · checked 10 Oct 2026
- 22agibot-world.com GO-1 page text (site bundle)Official site · 2025 · checked 10 Oct 2026
- 23AgiBot World Colosseo, full text v1Paper · Mar 2025 · checked 10 Oct 2026
- 24智元机器人开源百万真机数据集 AgiBot World (news report of the launch)Secondary · 30 Dec 2024 · checked 10 Oct 2026
- 25AgiBot World Colosseo project pageOfficial site · 10 Mar 2025 · checked 10 Oct 2026
- 26AgiBot World Colosseo (arXiv abstract page, submission history v1 to v4)Paper · Mar 2025 · checked 10 Oct 2026
- 27AgiBot-World issue #151: Alpha a subset of Beta? (maintainer reply)Repository · May 2026 · checked 10 Oct 2026
- 28Hugging Face Hub API: agibot-world/AgiBotWorld-Beta (downloads, likes, gated, lastModified)Index · 10 Oct 2026 · checked 10 Oct 2026
- 29Hugging Face Hub API: agibot-world/AgiBotWorld-AlphaIndex · 10 Oct 2026 · checked 10 Oct 2026
- 30agibot-world.com main site bundle (dataset, AGIBOT WORLD 2026 and Genie Sim text)Official site · 2026 · checked 10 Oct 2026
- 31OpenDriveLab/AgiBot-World commit historyRepository · 29 May 2026 · checked 10 Oct 2026
- 32Open-sourcing GO-1: The Bitter Lessons of Building VLA Systems at Scale (blog)Official site · 19 Sep 2025 · checked 10 Oct 2026
- 33Crossref record: AgiBot World Colosseo, IROS 2025, DOI 10.1109/IROS60139.2025.11247088Index · Oct 2025 · checked 10 Oct 2026
Where we searched for missing information
validity (sim-to-real or cross-site comparisons of AgiBot World scores): Paper v1 and v4 full text and figures; README; Hugging Face cards; OpenDriveLab project page and GO-1 blog; agibot-world.com page text (JS bundles for the dataset, GO-1, Genie Sim and both challenges); Genie Sim 3.0 paper v4 (its sim-real study uses other tasks); GitHub issues (160 issues and pull requests, titles read, key threads opened). No paired comparison involving AgiBot World's own evaluation tasks was found. Scores already come from real robots.
evaluation rubric and leaderboard: Paper Section V-A, README, both Hugging Face cards, project page, agibot-world.com bundles, GitHub evaluate/ folder (deploy.py, openloop_eval.py, LIBERO and AgileX examples). No per-task partial-credit rubric and no leaderboard for the dataset.
license_code: GitHub contents listing (no LICENSE file), GitHub licence API (404), pyproject.toml, README licence section.
published_at (IEEE TRO 2026 and best-paper-finalist claims): Crossref search for the paper title (only the IROS 2025 record) and for the follow-up title (TRO 2026 record); Semantic Scholar record of the follow-up (venue IEEE Transactions on Robotics). The IROS award claim was not checked: the shared web-search budget was used up before an IROS source could be found.
official Chinese launch announcement: agibot-world.com bundles (English and Chinese strings). A Chinese news report (IT之家, 2024-12-30) was read; thepaper.cn returned HTTP 403. No AgiBot press page was opened.
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).
- Expanded to a full entry from primary sources: paper v1 and v4 incl. figures, README, Hugging Face cards and API, GitHub issues, AgiBot site text, Crossref, Semantic Scholar and papers that reuse the data. Added the episode-versus-trajectory definition, size conflicts, data-quality issues, the partial-credit metric, paper results, challenges and reuse. Corrected: 'IEEE TRO 2026' belongs to the follow-up paper; scoring set to progress only.