CALVIN
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
CALVIN is a simulated benchmark in which a robot arm at a desk follows chains of five spoken-style instructions. It tests how many of the five tasks a policy (the robot's control model) completes in a row.12
What a score here does not tell you Inferred
- How well a policy will do on a real robot.We found no study that scored the same policies on CALVIN and on real robots.
- How well a policy copes with new start poses.When the block start poses were redrawn, scores fell by up to one task.
- Whether a policy understands new wording of an instruction.The test uses one fixed sentence for each subtask.
- Whether a small gain over earlier work is real.Fewer than half of the claimed gains can be shown to be statistically significant.
Comparisons with real robots
Tried on real robots, not compared Inferred1814+4
Details
The paper says CALVIN captures challenges of real-world settings but has no real-robot experiment. GR-1, MDT and FLOWER report real-robot results separately from CALVIN. PolaRiS cites CALVIN among simulation benchmarks that 'sacrifice realism' and measures it no further. The 2026 audit calls a sim-vs-real ranking test impractical and does not run one. The basic entry graded this 'claimed'; we raise it to 'demonstrated' because the taxonomy defines that level as 'some policies also ran on real robots'.
Our assessment Opinion
A high score shows that a policy can chain trained skills at one simulated desk. It does not show general skill.
Reasoning
A high CALVIN score shows that a policy can chain trained tabletop skills at this simulated desk. On ABC→D, it also shows that the policy copes with a new desk look and layout. On its own, the score is weak evidence of general manipulation skill. A small model given only a task number matches common policies, redrawn block poses cut scores by up to one task, and reworded instructions cut success by about 40 points. No study ties CALVIN scores to real-robot results.
Confidence: medium
Ignore gaps of about 0.1 between top results.
Reasoning
Treat gaps of about 0.1 in average length (the average number of tasks completed in a row) between top papers as ties. The same model is reported with differences of that size across versions of one paper, published rows do not always add up, and fewer than half of claimed ABC→D gains are provably significant.
Confidence: high
Check the split and the training data before comparing numbers.
Reasoning
Always check which split (the combination of training and test environments) a CALVIN number comes from. ABC→D tests an unseen environment. ABCD→D and D→D test an environment seen in training, and ABCD→D scores run higher. Also check the training data, because outside pretraining and RL (reinforcement learning, which trains by trial and error) in the test scene are allowed and change scores.
Confidence: high
The official leaderboard is out of date. Read the papers for current results.
Reasoning
The official leaderboard is a useful starting list, but it stops at September 2025 and differs from some source papers. Use the papers, or the 2026 audit's tracker, for current results.
Confidence: medium
Known problems 10
Fixes without version numbers changed the data and the evaluation
Bugs in the data and in the evaluation were fixed in 2022 and 2023 without version tags.222324+2
Details
The README changelog records a breaking change to evaluation start states (2022-01-10), changed success criteria for pushing and lifting (2022-02-07), wrong language annotations and scene files in the ABC and ABCD datasets fixed on 2022-09-16 (marked 'MAJOR BUG'), and a wrong scene file in the D dataset fixed on 2023-02-24. A bug in the LED button during rollouts was fixed in calvin_env on 2022-12-23 after issue #32. The repository has no version tags, so papers cannot state which version they used.
The dataset has no stated licence and one slow download source
The dataset has no stated licence. It is served from one HTTP server, and users report very slow downloads.27228+2
Details
No licence is stated for the CALVIN dataset. It is served from one university server over plain HTTP (ABC→D zip 555 GB). Users report downloads of tens of KB/s and broken unzips (issues #103 and #116, 2025-2026). Third-party copies on Hugging Face carry MIT, Apache-2.0, 'cc' or no licence labels.
Most claimed gains are not provably significant
Of 107 claimed improvements on ABC→D, 44% can be shown to be statistically significant.43132
Details
The 2026 audit sorted 107 previous-best-to-new comparisons on ABC→D, using only public scores. 47 (43.9%) are provably significant at the 5% level, 36 (33.6%) cannot be decided from averages, 21 (19.6%) show no improvement and 3 (2.8%) are provably not significant. The released CSV later moves those 3 to 'no improvement'.
A small model given only a task ID scores like widely used policies
A small model given only a task number, without the instruction, scored 3.24 on ABC→D. That is above RoboFlamingo but below the best results.43334+1
Details
The 2026 audit trained a 0.09B probe (DINOv2 encoder and an MLP) that receives a 34-way task ID instead of the instruction. It scored 3.242 on ABC→D, 3.872 on ABCD→D and 3.123 on D→D (best checkpoint; final checkpoints 3.149 and 3.783). That is above RoboFlamingo on ABC→D (2.48) and close to it on ABCD→D (4.09), but well below the best results (4.78 and 4.80). The audit concludes the shortcut on CALVIN is real but smaller than on LIBERO. The official evaluation uses one fixed sentence per subtask, so the sentence identifies the subtask.
Scores fall when block start poses are redrawn inside the training range
When the block start poses were redrawn inside the training range, the score of X-VLA fell from 4.17 to 3.14.43637
Details
The released ABC→D evaluation places blocks at fixed poses, while during training blocks start anywhere in a range. The 2026 audit redrew block poses from that same range and kept everything else fixed (1,000 chains). Average tasks completed fell from 4.165 to 3.138 for X-VLA (95% CI of the drop 0.890 to 1.164), from 3.244 to 2.495 for GR-1 and from 2.367 to 1.869 for RoboFlamingo. Fully completed chains fell by 25.0, 13.5 and 6.4 points. Two fresh 1,000-chain manifests with the original pose rules moved scores by at most 0.11, within noise. The audit calls this distribution overfitting.
The same result appears with different numbers
The paper on 3D Diffuser Actor reports its score as 3.27, 3.35 and 3.83. Some leaderboard rows do not add up.91038+7
Details
3D Diffuser Actor reports 3.27 (v1, 60 keyposes), 3.83 (v1, 360 keyposes) and 3.35 ± 0.04 (v3) on ABC→D; the leaderboard uses 3.27. VPP reports 4.29 (v1) and 4.33 (v2). RoboFlamingo reports 2.48 and 4.09; the leaderboard shows 2.47 and 4.08. DeeR-VLA is 2.82 on the leaderboard and 2.90 in FLOWER's table. FLOWER's own rows do not add up: its five in-a-row rates sum to 4.49, 4.74 and 4.33, against reported averages of 4.53, 4.67 and 4.35 (our arithmetic); the leaderboard copies these rows. The leaderboard misses every result after 2025-09. The 2026 audit's Table 1 lists MDT-V 4.52 as the best D→D score, but MDT reports 4.52 for ABCD→D and 3.72 for D→D. Two different models are both called UniVLA (3.80 and 4.41 on ABC→D).
Papers use different training recipes
Papers differ in their training data, in how many random seeds they average and in how they choose checkpoints (saved versions of a model).101216+3
Details
Papers train on different data for the same split: all play data or only the language-labelled windows (3D Diffuser Actor's table marks this per method), with or without outside pretraining such as internet video, DROID or Open X-Embodiment. Some average 3 seeds, others report one run; Seer averages its top 3 checkpoints, and the benchmark has no separate validation set. πRL adds RL in scene D and reports the result next to ABC→D numbers. ABC→D, ABCD→D and D→D are often reported without saying which training data were used.
Test runs differ across hardware
Changing only the GPU made the actions of GR-1 differ from the first step of a test run.224
Details
The README warns that GPU (EGL) rendering changes textures slightly compared with CPU rendering. The 2026 audit found that changing only the GPU (RTX A6000 vs RTX 6000 Ada) made GR-1's actions diverge from step 0. Changing only the CPU left 50 of 50 official sequences bitwise identical.
Top scores are close to the maximum of 5
The best average is 4.78 out of 5 on ABC→D, where policies train in environments A, B and C and are tested in D. On ABCD→D, which trains in all four, the best is 4.80.181741+1
Details
On ABC→D, published averages reach 4.78 (MMaDA-VLA, 2026-03; 89.7% of chains fully completed), 4.75 (Xiaomi-Robotics-0) and 4.73 (VITA). On ABCD→D the best is 4.80 (Xiaomi-Robotics-0; 91.8% of chains fully completed). Several results since late 2025 lie within 0.1 of each other. The CALVIN-Para authors describe canonical CALVIN single-task performance as near-saturated (92.3% and 99.7% for two models).
Policies fail when the instructions are reworded
When the instructions were reworded, success on single tasks fell by about 40 points.42735
Details
CALVIN-Para paraphrases 15 base tasks into 1,935 instructions and tests them one task at a time with official ABCD→D weights. Success fell from 92.3% to 53.1% for RoboFlamingo and from 99.7% to 58.3% for FLOWER (5 seeds). The RoboFlamingo paper tested GPT-4-rewritten instructions on scene D: ABCD→D average length fell from 4.09 to 1.85 for RoboFlamingo and from 3.06 to 1.82 for HULC. CALVIN's own test uses one fixed sentence per subtask.
Details
About
- What it is
- Benchmark1
More
The paper presents CALVIN as a benchmark with an environment, a dataset and a challenge with fixed evaluation protocols.
- Built by
- University of Freiburg, University of Technology Nuremberg12
More
Authors: Oier Mees and Lukas Hermann (equal contribution), Erick Rosete-Beas, Wolfram Burgard. Funded by the German Federal Ministry of Education and Research (contract 01IS18040B-OML).
University of Freiburg · Autonomous Intelligent Systems Lab. Affiliation of Oier Mees, Lukas Hermann and Erick Rosete-Beas.12
University of Technology Nuremberg · Affiliation of Wolfram Burgard on the paper.1
- Released
- December 2021. Published in RA-L in 2022.43144+1
More
arXiv v1 on 2021-12-06. Published in IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 7327-7334 (July 2022).
arXiv v4 (2022-07-13) is the accepted version (received 2022-02-23, accepted 2022-05-22). The README states that CALVIN won the 2022 RA-L Best Paper Award. The GitHub repository was created on 2021-07-20 (GitHub API).
- Version
- No releases. Users run the main branch.264522
More
No tags or releases. Users run the main branch and the dataset currently on the server. The README changelog records breaking changes on 2022-01-10, 2022-02-07, 2022-09-16 and 2023-02-24.
GitHub API: 0 releases, 0 tags.
Three training/test splits · D→D (train and test in environment D), ABCD→D (train in all four, test in D) and ABC→D (train in A, B and C, test in the unseen environment D).127
ABC→D is the most used split · The 2026 audit calls ABC→D CALVIN's most commonly used protocol.4
- Last update
- September 2025. Only a leaderboard entry was added.45228+1
More
2025-09-08: FLOWER added to the README model list and the website leaderboard (website Last-Modified header 2025-09-08). Last change to the evaluation code: 2023-12-07. Last dataset file change: the D→D zip, 2023-02-23.
Commit history: 2023-12-07 'fix small error in eval script at checkpoint loading'; 2024-02-08 visualisation fix; after that only README entries for new models. Server Last-Modified dates: task_D_D.zip 2023-02-23, task_ABC_D.zip and task_ABCD_D.zip 2022-09-15, debug zip 2022-05-13. The calvin_env submodule was last pushed 2024-01-03.
- Status
- No updates since September 2025. Use is very active. Inferred4523
More
No code change since 2024-02 and no leaderboard update since 2025-09-08. Use is very active.
By the taxonomy rule (no updates for over a year): last commit and website change 2025-09-08, about 13 months before 2026-10-10. 52 open issues. The basic entry said 'maintained'.
Setup
- Runs in
- Simulation1
More
All evaluation runs in simulation.
- Simulator
- PyBullet122
More
The README FAQ says EGL GPU rendering is used for speed and that textures render slightly differently on GPU than on CPU.
- Robot
- One arm1
- Robot model
- Franka Emika Panda (simulated, 7-DOF)146
More
7-DOF arm with a parallel gripper whose fingers cannot be controlled independently. Control at 30 Hz.
- Setting
- Tabletop1
- Tasks
- 34 tasks12
More
34 tasks, chained into sequences of 5
CONFLICT: the paper lists 34 tasks with success criteria (Fig. 6 and Fig. 9). The website leaderboard header says '(32 tasks)' for the MTLC column.
- Scenes
- 4 environments1
More
4 environments (A, B, C, D) with different textures and different positions of the static elements
- Training data
- About 24 hours of play data, with 1% labelled with language12747+1
More
About 24 hours of teleoperated play (about 6 hours per environment, about 2.4M interaction steps), collected by 3 untrained users with an HTC Vive VR headset. 1% of the data is labelled with language.
Play data has no fixed task list; users explored freely. Observations include static and gripper RGB-D cameras, tactile images and proprioception.
Language labels · Labelled automatically by a task detector. The appendix gives 389 unique instructions for 34 tasks (about 11 per task); the main text says 'over 400'. The introduction says 20K language directives.1
What '1%' means · The maintainers say 1% of 64-frame windows were labelled; training then cuts shorter sub-windows from them. A user counted about 40% of D→D training windows with language.47
Download sizes · D→D 177,379,436,142 bytes; ABC→D 555,309,812,705 bytes; ABCD→D 704,022,347,117 bytes; debug 1,299,150,917 bytes. The README gives 166, 517 and 656 GB, which match these sizes in GiB.2827
Precomputed language embeddings · MiniLM embeddings ship with the data; since 2022-09-16 nine more embedding sets are on the server.127
- Changes at test
- New scene look and layout, and new wording1374
More
ABC→D tests an unseen environment with different desk textures and moved drawer, sliding door, button and switch. Test instructions are phrasings not in the training set. D→D and ABCD→D test in an environment seen in training.
Block start poses are not varied much at test time. The evaluation code places blocks at two fixed table positions per start condition, with a rotation drawn under a fixed seed; the audit says the released scene-D evaluation fixes block poses. The paper says all scene elements of D appear, in other positions, in the training environments.
Scoring and access
- Scored by
- Chain length, Success rate1234
More
A subtask counts as solved when the simulator's task detector sees the required state change (for example a block lifted at least 5 cm).
- Score
- Tasks done in a row, out of 51234+1
More
Main metric (LH-MTLC): a policy gets 1,000 fixed chains of 5 instructions and moves to the next instruction only if it solved the current one. Scores are the share of chains with 1, 2, 3, 4 and 5 tasks done in a row, and their sum, the average number completed (Avg. Len.). A second metric (MTLC) scores single tasks, but the README says it is only available for the baseline agent.
The original paper reports only the five in-a-row rates; the website and later papers add Avg. Len. By arithmetic, Avg. Len. equals the sum of the five rates.
- Trials
- 1,000 chains of 5 tasks each13448+2
More
Official manifest · 1,000 chains generated with fixed random seeds from symbolic start states (drawer, slider, LED, light bulb and block locations). The robot is reset to a neutral pose before each chain.148
Time limit · 360 simulation steps per subtask (EP_LEN = 360), at 30 Hz.341
One test sentence per subtask · The evaluation script uses the first entry of the validation annotation file for each subtask; the file holds one instruction per task (34 lines).3435
Seeds · Some papers average 3 training seeds (HULC, 3D Diffuser Actor, MDT, MoDE, FLOWER); others report one run. Seer reports the average of its top 3 checkpoints.61019+2
- Who runs it
- Each team tests its own model222
More
The README asks authors to contact Oier Mees to add a model. The leaderboard copies numbers from papers; no organiser re-runs submissions.
- Error bars
- Sometimes reported Inferred2610+4
More
The website leaderboard shows point values only. HULC, 3D Diffuser Actor (v3), MDT, MoDE, GR-MG and FLOWER report standard deviations over 3 seeds. GR-1, RoboFlamingo, Seer, DreamVLA, X-VLA, Xiaomi-Robotics-0 and MMaDA-VLA report single numbers.
- Leaderboard
- Official and curated. Last updated in September 2025.2223
More
Maintainer-curated tables for D→D, ABCD→D and ABC→D on the project website. Last updated 2025-09-08.
Not comprehensive. It lists 46 rows from papers up to FLOWER (2025-09). Higher published results (for example 4.78 on ABC→D and 4.80 on ABCD→D) are missing. Several rows differ from the source papers (see issues.i5).
- Code licence
- MIT52
More
LICENSE: MIT, Copyright (c) 2021 Oier Mees. The website says the code is 'for academic usage and is released under the MIT license'; the MIT text itself has no academic-only condition.
- Data licence
- Unknown30
More
No licence for the dataset in the dataset README, download_data.sh, the main README or the website (all opened 2026-10-10). The dataset is served from calvin.cs.uni-freiburg.de. Third-party copies on Hugging Face carry self-assigned labels (items); they do not set CALVIN's own data licence.
MIT · InternRobotics/InternData-Calvin_ABC (third-party copy, 3,883 downloads)30
MIT · fywang/calvin-task-ABCD-D-lerobot (third-party copy)30
Apache-2.0 · ducido/calvin_task_D_D_* copies (third-party)30
none · zhouhongyi/calvin_abc (third-party copy with no licence tag; 44,088 downloads)30
- Asset licence
- MIT, with the robot model under Apache-2.0 Inferred4946
More
The simulation assets (desk, blocks, Franka Panda model) sit in the calvin_env repository under its MIT licence; the Franka Panda folder carries its own Apache License 2.0 file.
Read by us from the calvin_env repository tree: the only licence files are the root MIT LICENSE and data/franka_panda/LICENSE.txt (Apache 2.0). The origin of the desk textures is not stated. The tactile simulator is a fork of facebookresearch/tacto (MIT). Not legal advice.
- Access
- Open. HTTP download of 177 to 704 GB.272829
More
Code on GitHub. Data from the University of Freiburg server over plain HTTP, with SHA-256 checksums. No registration.
HTTPS to calvin.cs.uni-freiburg.de is refused (checked 2026-10-10). Users in issues #103 and #116 (2025) report download speeds of tens of KB/s and broken unzips; some use parallel downloaders or proxies. Evaluating any split needs the scene-D data.
Sources 49
- 1CALVIN paper, full text v4 (RA-L accepted version)Paper · Jul 2022 · checked 10 Oct 2026
- 2CALVIN project website and leaderboard (served over HTTP only; HTTPS refused)Leaderboard · 8 Sep 2025 · checked 10 Oct 2026
- 3Audit CALVIN citation tracker (calvin_citation_tracker.csv, snapshot 2026-05-21)Repository · May 2026 · checked 10 Oct 2026
- 4What Are We Actually Benchmarking in Robot Manipulation? (2026 audit; full text v1)Paper · Jun 2026 · checked 10 Oct 2026
- 5CALVIN LICENSE fileRepository · 2021 · checked 10 Oct 2026
- 6What Matters in Language Conditioned Robotic Imitation Learning over Unstructured Data (HULC)Paper · Apr 2022 · checked 10 Oct 2026
- 7Vision-Language Foundation Models as Effective Robot Imitators (RoboFlamingo)Paper · Nov 2023 · checked 10 Oct 2026
- 8Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (GR-1)Paper · Dec 2023 · checked 10 Oct 2026
- 93D Diffuser Actor: Policy Diffusion with 3D Scene Representations, v1Paper · Feb 2024 · checked 10 Oct 2026
- 103D Diffuser Actor, v3 (current version)Paper · Jul 2024 · checked 10 Oct 2026
- 11GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned PolicyPaper · Aug 2024 · checked 10 Oct 2026
- 12Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation (Seer)Paper · Dec 2024 · checked 10 Oct 2026
- 13DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgePaper · Jul 2025 · checked 10 Oct 2026
- 14FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow PoliciesPaper · Sep 2025 · checked 10 Oct 2026
- 15X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelPaper · Oct 2025 · checked 10 Oct 2026
- 16πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models (v3, Appendix C.3 and D.2)Paper · Oct 2025 · checked 10 Oct 2026
- 17Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution (v2)Paper · Feb 2026 · checked 10 Oct 2026
- 18MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation (v3)Paper · Mar 2026 · checked 10 Oct 2026
- 19Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals (MDT)Paper · Jul 2024 · checked 10 Oct 2026
- 20PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot PoliciesPaper · Dec 2025 · checked 10 Oct 2026
- 21A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA EvaluationPaper · Jun 2026 · checked 10 Oct 2026
- 22CALVIN GitHub README (evaluation, FAQ, changelog, model list)Repository · Sep 2025 · checked 10 Oct 2026
- 23calvin_env commit 797142c: fix bug in button during rolloutsRepository · 23 Dec 2022 · checked 11 Oct 2026
- 24CALVIN issue #32: Major concern about evaluationRepository · Dec 2022 · checked 11 Oct 2026
- 25CALVIN issue #40: some inconsistencies in the datasetRepository · Feb 2023 · checked 10 Oct 2026
- 26GitHub API: mees/calvin (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 27CALVIN dataset README (splits, sizes, data structure)Repository · 2022 · checked 10 Oct 2026
- 28CALVIN dataset server: sha256sum.txt and HTTP headers of the four zip files (size, Last-Modified)Repository · 24 Feb 2023 · checked 10 Oct 2026
- 29CALVIN issues #103 and #116: slow dataset downloads and broken unzipsRepository · Aug 2025 · checked 11 Oct 2026
- 30Hugging Face Hub dataset listing for 'calvin' (licence tags of third-party copies)Index · 10 Oct 2026 · checked 10 Oct 2026
- 31Audit release: statistical_significance_pie_counts.csv and statistical_significance_bucket_comparison.csvRepository · Oct 2026 · checked 10 Oct 2026
- 32Manipulation Benchmark Audit project pageOfficial site · 2026 · checked 10 Oct 2026
- 33Audit CALVIN probe results (shortcut_solvability/results/calvin summaries)Repository · Jun 2026 · checked 10 Oct 2026
- 34CALVIN evaluation script evaluate_policy.py (EP_LEN = 360, NUM_SEQUENCES = 1000, first validation instruction per subtask)Repository · 2023 · checked 10 Oct 2026
- 35CALVIN test instructions new_playtable_validation.yaml (one instruction per task)Repository · 2022 · checked 10 Oct 2026
- 36Audit CALVIN resampled-pose config and distribution_overfitting_summary.csvRepository · Jun 2026 · checked 10 Oct 2026
- 37CALVIN evaluation utils.py (get_env_state_for_initial_condition: fixed block positions)Repository · 2022 · checked 10 Oct 2026
- 38Video Prediction Policy (VPP), v1 and v2Paper · Dec 2024 · checked 10 Oct 2026
- 39Unified Vision-Language-Action Model (UniVLA, Wang et al.)Paper · Jun 2025 · checked 10 Oct 2026
- 40Learning to Act Anywhere with Task-centric Latent Actions (UniVLA, Bu et al.)Paper · May 2025 · checked 10 Oct 2026
- 41Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought (VITA)Paper · Nov 2025 · checked 10 Oct 2026
- 42LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models (v3, Appendix E: CALVIN-Para)Paper · Mar 2026 · checked 10 Oct 2026
- 43CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv abstract page, submission history v1 to v4)Paper · Dec 2021 · checked 10 Oct 2026
- 44Crossref record for DOI 10.1109/LRA.2022.3180108 (RA-L vol. 7, no. 3, pp. 7327-7334)Index · Jul 2022 · checked 11 Oct 2026
- 45CALVIN commit history (releases and tags lists are empty)Repository · 8 Sep 2025 · checked 10 Oct 2026
- 46calvin_env data/franka_panda/LICENSE.txt (Apache License 2.0)Repository · 2021 · checked 10 Oct 2026
- 47CALVIN issue #23: proportion of data with language instructionsRepository · Aug 2022 · checked 11 Oct 2026
- 48CALVIN sequence generator multistep_sequences.py (fixed seeds)Repository · 2022 · checked 10 Oct 2026
- 49calvin_env LICENSE (MIT) and repository tree (table, block and robot assets)Repository · 2021 · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: CALVIN paper v4 (no real-robot experiments); README and website; GR-1, MDT and FLOWER papers (real-robot results reported separately); PolaRiS (2512.16881: cites CALVIN as sacrificing realism, no measurement); 'A Practical Recipe Towards Improving Sim-and-Real Correlation' (2606.10366: CALVIN in related work only); X2Real (2609.27449: related work only); 2026 audit (calls the test impractical). One web search on 2026-10-10 for CALVIN sim-real correlation found no paired study.
license_data: dataset/README.md, download_data.sh, main README, LICENSE, website, calvin_env repository tree, dataset server directory (HTTP 403).
top_score (newer than 2026-05): 2026 audit tracker (snapshot 2026-05-21); two web searches on 2026-10-10 for CALVIN ABC→D averages above 4.78; arXiv API listing attempts on 2026-10-10 returned HTTP 503; the shared web-search budget then ran out. Not re-searched after that.
derived_benchmarks: One web search for CALVIN-based variants (2026-10-10), the LIBERO-Para, RoboFlamingo, GEVRM and audit papers.
leaderboard comprehensiveness: Website tables (46 rows) compared with the audit tracker and the source papers listed under top_score.
Change history
- Created at full depth from the checked basic entry and primary sources. Corrections to the basic entry: status set to dormant (no update for over a year), sim_to_real raised from 'claimed' to 'demonstrated', top score updated from the leaderboard (4.53) to the literature (4.78 ABC→D, 4.80 ABCD→D), leaderboard rows found inconsistent with source papers. Checking continued into 2026-10-11 (local time); sources opened then carry that access date.
- Published as a full entry.