CALVIN

CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

How to read this picture

CALVIN is a simulated benchmark in which a robot arm at a desk follows chains of five spoken-style instructions. It tests how many of the five tasks a policy (the robot's control model) completes in a row.12

Sources
Last checked 10 Oct 2026Full entry73 of 81 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation1
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm1
Used by
190 papers34
828 citations
Licence
MIT52
Commercial use: unclear

What a score here does not tell you Inferred

  1. How well a policy will do on a real robot.We found no study that scored the same policies on CALVIN and on real robots.
  2. How well a policy copes with new start poses.When the block start poses were redrawn, scores fell by up to one task.
  3. Whether a policy understands new wording of an instruction.The test uses one fixed sentence for each subtask.
  4. Whether a small gain over earlier work is real.Fewer than half of the claimed gains can be shown to be statistically significant.
ChartPublished scores over time
TOP 5% OF THE SCALE00.511.522.533.544.552023202420252026HULC · 0.67 · 2022-04RoboFlamingo · 2.48 · 2023-11GR-1 · 3.06 · 2023-123D Diffuser Actor · 3.27 · 2024-02GR-MG · 4.04 · 2024-08Seer-Large · 4.28 · 2024-12DreamVLA · 4.44 · 2025-07FLOWER · 4.53 · 2025-09X-VLA (0.9B) · 4.43 · 2025-10πRL on π0.5 (Flow-SDE) · 4.717 · 2025-10Xiaomi-Robotics-0 · 4.75 · 2026-02MMaDA-VLA · 4.78 · 2026-03HULC 0.67MMaDA-VLA 4.78
Score: Average number of tasks completed in a row, out of 5. Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.678+11

Comparisons with real robots

Tried on real robots, not compared Inferred1814+4

Details

The paper says CALVIN captures challenges of real-world settings but has no real-robot experiment. GR-1, MDT and FLOWER report real-robot results separately from CALVIN. PolaRiS cites CALVIN among simulation benchmarks that 'sacrifice realism' and measures it no further. The 2026 audit calls a sim-vs-real ranking test impractical and does not run one. The basic entry graded this 'claimed'; we raise it to 'demonstrated' because the taxonomy defines that level as 'some policies also ran on real robots'.

Our assessment Opinion

A high score shows that a policy can chain trained skills at one simulated desk. It does not show general skill.

Reasoning

A high CALVIN score shows that a policy can chain trained tabletop skills at this simulated desk. On ABC→D, it also shows that the policy copes with a new desk look and layout. On its own, the score is weak evidence of general manipulation skill. A small model given only a task number matches common policies, redrawn block poses cut scores by up to one task, and reworded instructions cut success by about 40 points. No study ties CALVIN scores to real-robot results.

Confidence: medium

Ignore gaps of about 0.1 between top results.

Reasoning

Treat gaps of about 0.1 in average length (the average number of tasks completed in a row) between top papers as ties. The same model is reported with differences of that size across versions of one paper, published rows do not always add up, and fewer than half of claimed ABC→D gains are provably significant.

Confidence: high

Check the split and the training data before comparing numbers.

Reasoning

Always check which split (the combination of training and test environments) a CALVIN number comes from. ABC→D tests an unseen environment. ABCD→D and D→D test an environment seen in training, and ABCD→D scores run higher. Also check the training data, because outside pretraining and RL (reinforcement learning, which trains by trial and error) in the test scene are allowed and change scores.

Confidence: high

The official leaderboard is out of date. Read the papers for current results.

Reasoning

The official leaderboard is a useful starting list, but it stops at September 2025 and differs from some source papers. Use the papers, or the 2026 audit's tracker, for current results.

Confidence: medium

Known problems 10

  1. Fixes without version numbers changed the data and the evaluation

    Bugs in the data and in the evaluation were fixed in 2022 and 2023 without version tags.222324+2

    Details

    The README changelog records a breaking change to evaluation start states (2022-01-10), changed success criteria for pushing and lifting (2022-02-07), wrong language annotations and scene files in the ABC and ABCD datasets fixed on 2022-09-16 (marked 'MAJOR BUG'), and a wrong scene file in the D dataset fixed on 2023-02-24. A bug in the LED button during rollouts was fixed in calvin_env on 2022-12-23 after issue #32. The repository has no version tags, so papers cannot state which version they used.

  2. The dataset has no stated licence and one slow download source

    The dataset has no stated licence. It is served from one HTTP server, and users report very slow downloads.27228+2

    Details

    No licence is stated for the CALVIN dataset. It is served from one university server over plain HTTP (ABC→D zip 555 GB). Users report downloads of tens of KB/s and broken unzips (issues #103 and #116, 2025-2026). Third-party copies on Hugging Face carry MIT, Apache-2.0, 'cc' or no licence labels.

  3. Most claimed gains are not provably significant

    Of 107 claimed improvements on ABC→D, 44% can be shown to be statistically significant.43132

    Details

    The 2026 audit sorted 107 previous-best-to-new comparisons on ABC→D, using only public scores. 47 (43.9%) are provably significant at the 5% level, 36 (33.6%) cannot be decided from averages, 21 (19.6%) show no improvement and 3 (2.8%) are provably not significant. The released CSV later moves those 3 to 'no improvement'.

  4. A small model given only a task ID scores like widely used policies

    A small model given only a task number, without the instruction, scored 3.24 on ABC→D. That is above RoboFlamingo but below the best results.43334+1

    Details

    The 2026 audit trained a 0.09B probe (DINOv2 encoder and an MLP) that receives a 34-way task ID instead of the instruction. It scored 3.242 on ABC→D, 3.872 on ABCD→D and 3.123 on D→D (best checkpoint; final checkpoints 3.149 and 3.783). That is above RoboFlamingo on ABC→D (2.48) and close to it on ABCD→D (4.09), but well below the best results (4.78 and 4.80). The audit concludes the shortcut on CALVIN is real but smaller than on LIBERO. The official evaluation uses one fixed sentence per subtask, so the sentence identifies the subtask.

  5. Scores fall when block start poses are redrawn inside the training range

    When the block start poses were redrawn inside the training range, the score of X-VLA fell from 4.17 to 3.14.43637

    Details

    The released ABC→D evaluation places blocks at fixed poses, while during training blocks start anywhere in a range. The 2026 audit redrew block poses from that same range and kept everything else fixed (1,000 chains). Average tasks completed fell from 4.165 to 3.138 for X-VLA (95% CI of the drop 0.890 to 1.164), from 3.244 to 2.495 for GR-1 and from 2.367 to 1.869 for RoboFlamingo. Fully completed chains fell by 25.0, 13.5 and 6.4 points. Two fresh 1,000-chain manifests with the original pose rules moved scores by at most 0.11, within noise. The audit calls this distribution overfitting.

  6. The same result appears with different numbers

    The paper on 3D Diffuser Actor reports its score as 3.27, 3.35 and 3.83. Some leaderboard rows do not add up.91038+7

    Details

    3D Diffuser Actor reports 3.27 (v1, 60 keyposes), 3.83 (v1, 360 keyposes) and 3.35 ± 0.04 (v3) on ABC→D; the leaderboard uses 3.27. VPP reports 4.29 (v1) and 4.33 (v2). RoboFlamingo reports 2.48 and 4.09; the leaderboard shows 2.47 and 4.08. DeeR-VLA is 2.82 on the leaderboard and 2.90 in FLOWER's table. FLOWER's own rows do not add up: its five in-a-row rates sum to 4.49, 4.74 and 4.33, against reported averages of 4.53, 4.67 and 4.35 (our arithmetic); the leaderboard copies these rows. The leaderboard misses every result after 2025-09. The 2026 audit's Table 1 lists MDT-V 4.52 as the best D→D score, but MDT reports 4.52 for ABCD→D and 3.72 for D→D. Two different models are both called UniVLA (3.80 and 4.41 on ABC→D).

  7. Papers use different training recipes

    Papers differ in their training data, in how many random seeds they average and in how they choose checkpoints (saved versions of a model).101216+3

    Details

    Papers train on different data for the same split: all play data or only the language-labelled windows (3D Diffuser Actor's table marks this per method), with or without outside pretraining such as internet video, DROID or Open X-Embodiment. Some average 3 seeds, others report one run; Seer averages its top 3 checkpoints, and the benchmark has no separate validation set. πRL adds RL in scene D and reports the result next to ABC→D numbers. ABC→D, ABCD→D and D→D are often reported without saying which training data were used.

  8. Test runs differ across hardware

    Changing only the GPU made the actions of GR-1 differ from the first step of a test run.224

    Details

    The README warns that GPU (EGL) rendering changes textures slightly compared with CPU rendering. The 2026 audit found that changing only the GPU (RTX A6000 vs RTX 6000 Ada) made GR-1's actions diverge from step 0. Changing only the CPU left 50 of 50 official sequences bitwise identical.

  9. Top scores are close to the maximum of 5

    The best average is 4.78 out of 5 on ABC→D, where policies train in environments A, B and C and are tested in D. On ABCD→D, which trains in all four, the best is 4.80.181741+1

    Details

    On ABC→D, published averages reach 4.78 (MMaDA-VLA, 2026-03; 89.7% of chains fully completed), 4.75 (Xiaomi-Robotics-0) and 4.73 (VITA). On ABCD→D the best is 4.80 (Xiaomi-Robotics-0; 91.8% of chains fully completed). Several results since late 2025 lie within 0.1 of each other. The CALVIN-Para authors describe canonical CALVIN single-task performance as near-saturated (92.3% and 99.7% for two models).

  10. Policies fail when the instructions are reworded

    When the instructions were reworded, success on single tasks fell by about 40 points.42735

    Details

    CALVIN-Para paraphrases 15 base tasks into 1,935 instructions and tests them one task at a time with official ABCD→D weights. Success fell from 92.3% to 53.1% for RoboFlamingo and from 99.7% to 58.3% for FLOWER (5 seeds). The RoboFlamingo paper tested GPT-4-rewritten instructions on scene D: ABCD→D average length fell from 4.09 to 1.85 for RoboFlamingo and from 3.06 to 1.82 for HULC. CALVIN's own test uses one fixed sentence per subtask.

Details

About

What it is
Benchmark1
More

The paper presents CALVIN as a benchmark with an environment, a dataset and a challenge with fixed evaluation protocols.

Built by
University of Freiburg, University of Technology Nuremberg12
More

Authors: Oier Mees and Lukas Hermann (equal contribution), Erick Rosete-Beas, Wolfram Burgard. Funded by the German Federal Ministry of Education and Research (contract 01IS18040B-OML).

University of Freiburg · Autonomous Intelligent Systems Lab. Affiliation of Oier Mees, Lukas Hermann and Erick Rosete-Beas.12

University of Technology Nuremberg · Affiliation of Wolfram Burgard on the paper.1

Released
December 2021. Published in RA-L in 2022.43144+1
More

arXiv v1 on 2021-12-06. Published in IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 7327-7334 (July 2022).

arXiv v4 (2022-07-13) is the accepted version (received 2022-02-23, accepted 2022-05-22). The README states that CALVIN won the 2022 RA-L Best Paper Award. The GitHub repository was created on 2021-07-20 (GitHub API).

Version
No releases. Users run the main branch.264522
More

No tags or releases. Users run the main branch and the dataset currently on the server. The README changelog records breaking changes on 2022-01-10, 2022-02-07, 2022-09-16 and 2023-02-24.

GitHub API: 0 releases, 0 tags.

Three training/test splits · D→D (train and test in environment D), ABCD→D (train in all four, test in D) and ABC→D (train in A, B and C, test in the unseen environment D).127

ABC→D is the most used split · The 2026 audit calls ABC→D CALVIN's most commonly used protocol.4

Last update
September 2025. Only a leaderboard entry was added.45228+1
More

2025-09-08: FLOWER added to the README model list and the website leaderboard (website Last-Modified header 2025-09-08). Last change to the evaluation code: 2023-12-07. Last dataset file change: the D→D zip, 2023-02-23.

Commit history: 2023-12-07 'fix small error in eval script at checkpoint loading'; 2024-02-08 visualisation fix; after that only README entries for new models. Server Last-Modified dates: task_D_D.zip 2023-02-23, task_ABC_D.zip and task_ABCD_D.zip 2022-09-15, debug zip 2022-05-13. The calvin_env submodule was last pushed 2024-01-03.

Status
No updates since September 2025. Use is very active. Inferred4523
More

No code change since 2024-02 and no leaderboard update since 2025-09-08. Use is very active.

By the taxonomy rule (no updates for over a year): last commit and website change 2025-09-08, about 13 months before 2026-10-10. 52 open issues. The basic entry said 'maintained'.

Setup

Runs in
Simulation1
More

All evaluation runs in simulation.

Simulator
PyBullet122
More

The README FAQ says EGL GPU rendering is used for speed and that textures render slightly differently on GPU than on CPU.

Robot
One arm1
Robot model
Franka Emika Panda (simulated, 7-DOF)146
More

7-DOF arm with a parallel gripper whose fingers cannot be controlled independently. Control at 30 Hz.

Setting
Tabletop1
Tasks
34 tasks12
More

34 tasks, chained into sequences of 5

CONFLICT: the paper lists 34 tasks with success criteria (Fig. 6 and Fig. 9). The website leaderboard header says '(32 tasks)' for the MTLC column.

Scenes
4 environments1
More

4 environments (A, B, C, D) with different textures and different positions of the static elements

Training data
About 24 hours of play data, with 1% labelled with language12747+1
More

About 24 hours of teleoperated play (about 6 hours per environment, about 2.4M interaction steps), collected by 3 untrained users with an HTC Vive VR headset. 1% of the data is labelled with language.

Play data has no fixed task list; users explored freely. Observations include static and gripper RGB-D cameras, tactile images and proprioception.

Language labels · Labelled automatically by a task detector. The appendix gives 389 unique instructions for 34 tasks (about 11 per task); the main text says 'over 400'. The introduction says 20K language directives.1

What '1%' means · The maintainers say 1% of 64-frame windows were labelled; training then cuts shorter sub-windows from them. A user counted about 40% of D→D training windows with language.47

Download sizes · D→D 177,379,436,142 bytes; ABC→D 555,309,812,705 bytes; ABCD→D 704,022,347,117 bytes; debug 1,299,150,917 bytes. The README gives 166, 517 and 656 GB, which match these sizes in GiB.2827

Precomputed language embeddings · MiniLM embeddings ship with the data; since 2022-09-16 nine more embedding sets are on the server.127

Changes at test
New scene look and layout, and new wording1374
More

ABC→D tests an unseen environment with different desk textures and moved drawer, sliding door, button and switch. Test instructions are phrasings not in the training set. D→D and ABCD→D test in an environment seen in training.

Block start poses are not varied much at test time. The evaluation code places blocks at two fixed table positions per start condition, with a rotation drawn under a fixed seed; the audit says the released scene-D evaluation fixes block poses. The paper says all scene elements of D appear, in other positions, in the training environments.

Scoring and access

Scored by
Chain length, Success rate1234
More

A subtask counts as solved when the simulator's task detector sees the required state change (for example a block lifted at least 5 cm).

Score
Tasks done in a row, out of 51234+1
More

Main metric (LH-MTLC): a policy gets 1,000 fixed chains of 5 instructions and moves to the next instruction only if it solved the current one. Scores are the share of chains with 1, 2, 3, 4 and 5 tasks done in a row, and their sum, the average number completed (Avg. Len.). A second metric (MTLC) scores single tasks, but the README says it is only available for the baseline agent.

The original paper reports only the five in-a-row rates; the website and later papers add Avg. Len. By arithmetic, Avg. Len. equals the sum of the five rates.

Trials
1,000 chains of 5 tasks each13448+2
More

Official manifest · 1,000 chains generated with fixed random seeds from symbolic start states (drawer, slider, LED, light bulb and block locations). The robot is reset to a neutral pose before each chain.148

Time limit · 360 simulation steps per subtask (EP_LEN = 360), at 30 Hz.341

One test sentence per subtask · The evaluation script uses the first entry of the validation annotation file for each subtask; the file holds one instruction per task (34 lines).3435

Seeds · Some papers average 3 training seeds (HULC, 3D Diffuser Actor, MDT, MoDE, FLOWER); others report one run. Seer reports the average of its top 3 checkpoints.61019+2

Who runs it
Each team tests its own model222
More

The README asks authors to contact Oier Mees to add a model. The leaderboard copies numbers from papers; no organiser re-runs submissions.

Error bars
Sometimes reported Inferred2610+4
More

The website leaderboard shows point values only. HULC, 3D Diffuser Actor (v3), MDT, MoDE, GR-MG and FLOWER report standard deviations over 3 seeds. GR-1, RoboFlamingo, Seer, DreamVLA, X-VLA, Xiaomi-Robotics-0 and MMaDA-VLA report single numbers.

Leaderboard
Official and curated. Last updated in September 2025.2223
More

Maintainer-curated tables for D→D, ABCD→D and ABC→D on the project website. Last updated 2025-09-08.

Not comprehensive. It lists 46 rows from papers up to FLOWER (2025-09). Higher published results (for example 4.78 on ABC→D and 4.80 on ABCD→D) are missing. Several rows differ from the source papers (see issues.i5).

Code licence
MIT52
More

LICENSE: MIT, Copyright (c) 2021 Oier Mees. The website says the code is 'for academic usage and is released under the MIT license'; the MIT text itself has no academic-only condition.

Data licence
Unknown30
More

No licence for the dataset in the dataset README, download_data.sh, the main README or the website (all opened 2026-10-10). The dataset is served from calvin.cs.uni-freiburg.de. Third-party copies on Hugging Face carry self-assigned labels (items); they do not set CALVIN's own data licence.

MIT · InternRobotics/InternData-Calvin_ABC (third-party copy, 3,883 downloads)30

MIT · fywang/calvin-task-ABCD-D-lerobot (third-party copy)30

Apache-2.0 · ducido/calvin_task_D_D_* copies (third-party)30

none · zhouhongyi/calvin_abc (third-party copy with no licence tag; 44,088 downloads)30

Asset licence
MIT, with the robot model under Apache-2.0 Inferred4946
More

The simulation assets (desk, blocks, Franka Panda model) sit in the calvin_env repository under its MIT licence; the Franka Panda folder carries its own Apache License 2.0 file.

Read by us from the calvin_env repository tree: the only licence files are the root MIT LICENSE and data/franka_panda/LICENSE.txt (Apache 2.0). The origin of the desk textures is not stated. The tactile simulator is a fork of facebookresearch/tacto (MIT). Not legal advice.

Access
Open. HTTP download of 177 to 704 GB.272829
More

Code on GitHub. Data from the University of Freiburg server over plain HTTP, with SHA-256 checksums. No registration.

HTTPS to calvin.cs.uni-freiburg.de is refused (checked 2026-10-10). Users in issues #103 and #116 (2025) report download speeds of tens of KB/s and broken unzips; some use parallel downloaders or proxies. Evaluating any split needs the scene-D data.

Commercial use
Unclear Inferred54946+2
More

Code (MIT) and simulation assets (MIT, Apache-2.0) allow commercial use with attribution. No licence is stated for the dataset, so its terms are unknown. The website's phrase 'for academic usage' conflicts in tone with the MIT licence text. Not legal advice.

Published at
RA-L 2022441
More

DOI 10.1109/LRA.2022.3180108.

Sources 49

  1. 1CALVIN paper, full text v4 (RA-L accepted version)Paper · Jul 2022 · checked 10 Oct 2026
  2. 2CALVIN project website and leaderboard (served over HTTP only; HTTPS refused)Leaderboard · 8 Sep 2025 · checked 10 Oct 2026
  3. 3Audit CALVIN citation tracker (calvin_citation_tracker.csv, snapshot 2026-05-21)Repository · May 2026 · checked 10 Oct 2026
  4. 4What Are We Actually Benchmarking in Robot Manipulation? (2026 audit; full text v1)Paper · Jun 2026 · checked 10 Oct 2026
  5. 5CALVIN LICENSE fileRepository · 2021 · checked 10 Oct 2026
  6. 6What Matters in Language Conditioned Robotic Imitation Learning over Unstructured Data (HULC)Paper · Apr 2022 · checked 10 Oct 2026
  7. 7Vision-Language Foundation Models as Effective Robot Imitators (RoboFlamingo)Paper · Nov 2023 · checked 10 Oct 2026
  8. 8Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (GR-1)Paper · Dec 2023 · checked 10 Oct 2026
  9. 93D Diffuser Actor: Policy Diffusion with 3D Scene Representations, v1Paper · Feb 2024 · checked 10 Oct 2026
  10. 103D Diffuser Actor, v3 (current version)Paper · Jul 2024 · checked 10 Oct 2026
  11. 11GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned PolicyPaper · Aug 2024 · checked 10 Oct 2026
  12. 12Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation (Seer)Paper · Dec 2024 · checked 10 Oct 2026
  13. 13DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgePaper · Jul 2025 · checked 10 Oct 2026
  14. 14FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow PoliciesPaper · Sep 2025 · checked 10 Oct 2026
  15. 15X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelPaper · Oct 2025 · checked 10 Oct 2026
  16. 16πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (v3, Appendix C.3 and D.2)Paper · Oct 2025 · checked 10 Oct 2026
  17. 17Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution (v2)Paper · Feb 2026 · checked 10 Oct 2026
  18. 18MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation (v3)Paper · Mar 2026 · checked 10 Oct 2026
  19. 19Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals (MDT)Paper · Jul 2024 · checked 10 Oct 2026
  20. 20PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
  21. 21A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA EvaluationPaper · Jun 2026 · checked 10 Oct 2026
  22. 22CALVIN GitHub README (evaluation, FAQ, changelog, model list)Repository · Sep 2025 · checked 10 Oct 2026
  23. 23calvin_env commit 797142c: fix bug in button during rolloutsRepository · 23 Dec 2022 · checked 11 Oct 2026
  24. 24CALVIN issue #32: Major concern about evaluationRepository · Dec 2022 · checked 11 Oct 2026
  25. 25CALVIN issue #40: some inconsistencies in the datasetRepository · Feb 2023 · checked 10 Oct 2026
  26. 26GitHub API: mees/calvin (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
  27. 27CALVIN dataset README (splits, sizes, data structure)Repository · 2022 · checked 10 Oct 2026
  28. 28CALVIN dataset server: sha256sum.txt and HTTP headers of the four zip files (size, Last-Modified)Repository · 24 Feb 2023 · checked 10 Oct 2026
  29. 29CALVIN issues #103 and #116: slow dataset downloads and broken unzipsRepository · Aug 2025 · checked 11 Oct 2026
  30. 30Hugging Face Hub dataset listing for 'calvin' (licence tags of third-party copies)Index · 10 Oct 2026 · checked 10 Oct 2026
  31. 31Audit release: statistical_significance_pie_counts.csv and statistical_significance_bucket_comparison.csvRepository · Oct 2026 · checked 10 Oct 2026
  32. 32Manipulation Benchmark Audit project pageOfficial site · 2026 · checked 10 Oct 2026
  33. 33Audit CALVIN probe results (shortcut_solvability/results/calvin summaries)Repository · Jun 2026 · checked 10 Oct 2026
  34. 34CALVIN evaluation script evaluate_policy.py (EP_LEN = 360, NUM_SEQUENCES = 1000, first validation instruction per subtask)Repository · 2023 · checked 10 Oct 2026
  35. 35CALVIN test instructions new_playtable_validation.yaml (one instruction per task)Repository · 2022 · checked 10 Oct 2026
  36. 36Audit CALVIN resampled-pose config and distribution_overfitting_summary.csvRepository · Jun 2026 · checked 10 Oct 2026
  37. 37CALVIN evaluation utils.py (get_env_state_for_initial_condition: fixed block positions)Repository · 2022 · checked 10 Oct 2026
  38. 38Video Prediction Policy (VPP), v1 and v2Paper · Dec 2024 · checked 10 Oct 2026
  39. 39Unified Vision-Language-Action Model (UniVLA, Wang et al.)Paper · Jun 2025 · checked 10 Oct 2026
  40. 40Learning to Act Anywhere with Task-centric Latent Actions (UniVLA, Bu et al.)Paper · May 2025 · checked 10 Oct 2026
  41. 41Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought (VITA)Paper · Nov 2025 · checked 10 Oct 2026
  42. 42LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models (v3, Appendix E: CALVIN-Para)Paper · Mar 2026 · checked 10 Oct 2026
  43. 43CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv abstract page, submission history v1 to v4)Paper · Dec 2021 · checked 10 Oct 2026
  44. 44Crossref record for DOI 10.1109/LRA.2022.3180108 (RA-L vol. 7, no. 3, pp. 7327-7334)Index · Jul 2022 · checked 11 Oct 2026
  45. 45CALVIN commit history (releases and tags lists are empty)Repository · 8 Sep 2025 · checked 10 Oct 2026
  46. 46calvin_env data/franka_panda/LICENSE.txt (Apache License 2.0)Repository · 2021 · checked 10 Oct 2026
  47. 47CALVIN issue #23: proportion of data with language instructionsRepository · Aug 2022 · checked 11 Oct 2026
  48. 48CALVIN sequence generator multistep_sequences.py (fixed seeds)Repository · 2022 · checked 10 Oct 2026
  49. 49calvin_env LICENSE (MIT) and repository tree (table, block and robot assets)Repository · 2021 · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: CALVIN paper v4 (no real-robot experiments); README and website; GR-1, MDT and FLOWER papers (real-robot results reported separately); PolaRiS (2512.16881: cites CALVIN as sacrificing realism, no measurement); 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366: CALVIN in related work only); X2Real (2609.27449: related work only); 2026 audit (calls the test impractical). One web search on 2026-10-10 for CALVIN sim-real correlation found no paired study.

license_data: dataset/README.md, download_data.sh, main README, LICENSE, website, calvin_env repository tree, dataset server directory (HTTP 403).

top_score (newer than 2026-05): 2026 audit tracker (snapshot 2026-05-21); two web searches on 2026-10-10 for CALVIN ABC→D averages above 4.78; arXiv API listing attempts on 2026-10-10 returned HTTP 503; the shared web-search budget then ran out. Not re-searched after that.

derived_benchmarks: One web search for CALVIN-based variants (2026-10-10), the LIBERO-Para, RoboFlamingo, GEVRM and audit papers.

leaderboard comprehensiveness: Website tables (46 rows) compared with the audit tracker and the source papers listed under top_score.

Change history

  1. Created at full depth from the checked basic entry and primary sources. Corrections to the basic entry: status set to dormant (no update for over a year), sim_to_real raised from 'claimed' to 'demonstrated', top score updated from the leaderboard (4.53) to the literature (4.78 ABC→D, 4.80 ABCD→D), leaderboard rows found inconsistent with source papers. Checking continued into 2026-10-11 (local time); sources opened then carry that access date.
  2. Published as a full entry.