PolaRiS
PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies
PolaRiS builds simulated test scenes from short video scans of real scenes, to score robot policies (the robots' control models) for the DROID robot setup. Its authors report that the scores track real-robot results after each policy is briefly fine-tuned with some simulated data (co-training).12
What a score here does not tell you Inferred
- How a policy performs without changes.Each policy is scored after 1k steps of fine-tuning on simulated data.
- How a policy does on robots other than the DROID Franka arm.PolaRiS supports only the DROID setup with joint-position control.
- How a policy handles soft objects or tasks with complex contact.All six tasks use rigid objects. The authors say the tasks leave out non-rigid objects.
Comparisons with real robots
| Study | Result | What was compared | Done by |
|---|---|---|---|
| PolaRiS paper (arXiv) Dec 2025 | Mean over 6 scenes: Pearson r = 0.90, MMRV = 0.03. MMRV measures how often two rankings disagree.1 | 5 policies (pi0, pi0 at 100k steps, pi0-FAST, pi0.5 and PaliGemma-binning, so 4 models) were scored in simulation after 1k co-training steps, with 50 rollouts per task. They were also scored without co-training on the matching real scenes, with 20 rollouts per policy and scene. There were 6 scenes at UW and Princeton, scored by task progress. The study’s authors described the result as “strong”. | The benchmark’s authors |
| PolaRiS paper (RSS 2026 version) Jul 2026 | Pearson r = 0.84 and MMRV = 0.08 in Fig. 6, with 95% intervals. The text gives r = 0.83.2 | The same policies and scenes were used, with bootstrap resampling (repeated random sampling) of 50 episodes per policy, repeated 500 times. The study’s authors described the result as “strong”. | The benchmark’s authors |
| PolaRiS compared with RoboArena Dec 2025 | Pearson r = 0.98, MMRV = 0.00124 | The mean PolaRiS scores of 4 policies were compared with their mean progress scores in RoboArena's crowd-sourced real-robot evaluations. RoboArena uses different and broader tasks. The study’s authors described the result as “strong”. | The benchmark’s authors |
| PolaRiS paper: without co-training Dec 2025 | Pearson r = 0.30 and MMRV = 0.29. With co-training, r = 0.86 and MMRV = 0.07.1 | The same policies were evaluated zero-shot, with no co-training in simulation, on 3 target scenes with 5 seeds. The study’s authors described the result as “too low to accurately rank”. | The benchmark’s authors |
| SimFoundry (NVIDIA, Georgia Tech, Stanford, UT Austin, Toronto) Jun 2026 | Pearson r from -0.396 to 0.822 and MMRV from 0.053 to 0.352, depending on the task3 | 5 policies (pi0, pi0.5, GR00T N1.6, GR00T N1.7 and DreamZero) were tested zero-shot on 4 tasks, and 3 fine-tuned policies on 3 tasks. The tests used DROID scenes rebuilt with PolaRiS's tools, without co-training in simulation. Each policy ran 25 rollouts per task, both on the real robot and in simulation. The study’s authors described the result as “low correlation”. | An independent group |
Benchmarks built on PolaRiS: H2RBench.6
Our assessment Opinion
PolaRiS is useful for ranking DROID policies, but only with the co-training step.
Reasoning
Treat PolaRiS scores as estimates of how DROID-style vision-language-action (VLA) models rank on rigid tabletop tasks. These estimates are valid only under the co-training protocol. The supporting evidence is the authors' study of 5 closely related policies. The one independent test skipped co-training and found low correlation, which matches the authors' own result without co-training (zero-shot).
Confidence: medium
Use the UW scenes until the Princeton reset states are fixed.
Reasoning
Use the three UW scenes until the Princeton reset states are fixed. The author has confirmed that the Princeton reset states are broken, and an outside run scored far below the paper on two of those scenes.
Confidence: high
New scenes are cheap to build. The correlation must be checked again in each new scene.
Reasoning
PolaRiS's main value is that new test scenes are cheap to build (under 20 minutes of human time, under an hour in total), so teams can test in their own deployment scenes. Whether the correlation holds in a new scene has to be checked there.
Confidence: medium
Known problems 5, 1 disputed
Three of the six public environments do not reproduce the paper
The author confirmed broken reset states in 3 of the 6 public scenes.91011
Details
In March 2026 a user reported start states that made Move Latte Cup and Tape Into Container impossible. The author uploaded new initial conditions on 2026-03-14 without testing them; the user still saw objects stuck or falling off the table. In June 2026 another user ran the released pi0.5 checkpoint and got much lower scores than reported on the Princeton scenes: task progress 33.3% vs 60.0% reported on Organize Tools (0/100 successes), 34.3% vs 80.0% on Tape Into Container. On 2026-06-24 the author confirmed the Princeton reset states are broken, said the UW scenes should be fine, and said a fix will come with a larger release. The user also hit occasional crashes in the public code, which the author has seen too.
Scores depend on fine-tuning each policy in simulation
Each policy must first be fine-tuned on simulated data. Without this step, the correlation with real results was r = 0.30.12
Details
PolaRiS's protocol co-finetunes every policy for 1k steps on 90% real DROID data and 10% simulated data before scoring it. The policy scored in simulation is therefore not the one deployed on the real robot. In the authors' own ablation, the same policies evaluated without co-training gave Pearson r 0.30 and MMRV 0.29 (co-trained: 0.86 and 0.07). Fine-tuning for too long also lowers correlation. Policies whose weights cannot be fine-tuned cannot be evaluated this way (our inference).
Headline numbers differ between versions
The arXiv version reports a correlation of r = 0.90. The RSS version reports 0.83.1212
Details
arXiv v2 reports mean Pearson r 0.90 and MMRV 0.03 (the plain mean over six scenes). The RSS 2026 version reports r 0.83 in the text and 0.84 in Fig. 6, with MMRV 0.08, after bootstrap resampling. Comparison numbers changed too: Ctrl-World MMRV 0.22 to 0.23, LIBERO r 0.66 / 0.70 / 0.66 to 0.63 / 0.68 / 0.64. The released co-training data has 316 episodes, against 326 in the paper's table and about 350 in its text.
The validation set is small and narrow
The check against real robots used 5 related policies and 6 rigid-object tasks.12
Details
The real-robot check used 5 policies from one family (openpi DROID policies, with pi0 at two training stages), 6 rigid-object tasks at two institutions, 20 real rollouts per policy and scene, and human-scored progress. Each per-scene correlation rests on 5 points. The authors say the tasks leave out non-rigid objects, soft-body interaction and complex contact, and that PolaRiS does not replace real-world evaluation.
An independent test found low correlation without co-trainingDisputed
An outside group skipped co-training and got correlations from -0.396 to 0.822, depending on the task.3
Other view: The outside test skipped the co-training step that PolaRiS requires.31
Details
SimFoundry's authors rebuilt their own DROID scenes with PolaRiS's tools and scored 5 policies (pi0, pi0.5, GR00T N1.6, GR00T N1.7, DreamZero) zero-shot, 25 rollouts per policy and task, real and simulated. Per-task Pearson r ranged from -0.396 to 0.822 (undefined on one task), MMRV from 0.053 to 0.352, and most policies scored far below their real success.
PolaRiS's protocol requires co-training, which SimFoundry skipped for both systems; the PolaRiS authors' own zero-shot ablation also gave low correlation (r 0.30). PolaRiS's co-trained pi0.5 scored higher in SimFoundry's scenes but was left out of the correlation.
Details
About
- What it is
- Benchmark Inferred113
More
Ships 6 fixed evaluation environments with initial conditions and scoring rubrics, plus tools to build more. Classification by the Atlas.
- Built by
- University of Washington, Princeton University, UC Berkeley, Stanford University, Toyota Research Institute, University of Southern California, Cornell University, Physical Intelligence14115
More
14 authors. Equal contribution: Arhan Jain (UW), Mingtong Zhang (Princeton). Equal advising: Abhishek Gupta (UW, TRI), Karl Pertsch (Berkeley, Stanford, Physical Intelligence). Sergey Levine and Chelsea Finn also list Physical Intelligence.
- Released
- December 2025, at RSS 20261416
More
arXiv v1 on 2025-12-18. Published at RSS 2026 (Robotics: Science and Systems XXII, Sydney, 13-17 July 2026).
Code repository created 2025-11-24; Hub dataset created 2025-11-29.
- Version
- 0.1.0. There are no tagged releases.171814+1
More
Package polaris 0.1.0. No tags or releases. Paper: arXiv v1 (2025-12-18) and v2 (2025-12-30, references and acknowledgements only), then the RSS 2026 version with revised numbers.
arXiv v2 vs RSS 2026 · arXiv: mean r 0.90, MMRV 0.03, 'two institutions'. RSS: r 0.83 in the text (0.84 in Fig. 6), MMRV 0.08, bootstrap intervals, a new physics-randomisation test, 600 real and over 93,000 simulated rollouts.12
Hub data change 2026-03-14 · New initial conditions for the three Princeton environments (commit 'fix princeton ics'), no changelog. Copies downloaded earlier hold the old files.119
Setup
- Runs in
- Simulation1
More
Real-to-sim: reconstructions of real scenes.
- Simulator
- Isaac Sim, with Gaussian splat rendering (scenes rebuilt from video)117
More
2DGS meshes give collision geometry and splats give images; objects are generated from photos with TRELLIS; scenes are composed in a web GUI and exported as USD. Policies use joint-position actions; DROID control runs at 15 Hz (Appendix D).
- Robot
- One arm1
- Robot model
- DROID platform: Franka Panda 7-DoF arm, Robotiq 2F-85 gripper, one wrist and one external ZED camera119
More
Gripper model read from the Hub's robot folder file names. Supports wrist cameras, unlike SimplerEnv's green-screening.
- Setting
- Tabletop, Kitchen, Office or lab Inferred119
More
Tabletop tasks on a kitchen table, a stovetop, a corner table and lab benches (from scene asset names). 3 scenes at UW and 3 at Princeton.
- Tasks
- 6 tasks in 6 scenes113
More
6 evaluation tasks, one per scene: block stacking, food bussing, pan cleaning, move latte cup, organize tools, tape into container. 15 further scenes are used only for co-training.
- Training data
- 316 simulated demonstrations for co-training Inferred121
More
316 simulated teleoperated episodes in the released co-training dataset (RLDS, 3.1 GB, 15 scenes). The paper's table lists 326 trajectories (56,041 timesteps) for this data mix; its text says about 350.
CONFLICT. 316 is our sum of the shard lengths in dataset_info.json. Collected with a VR controller through the DROID teleoperation code. Each policy is co-finetuned on 10% of this data and 90% real DROID data for 1k steps.
- Changes at test
- Unseen scenes and varied start positions Inferred113
More
The 6 evaluation scenes are unseen by the sim co-training step; object start positions vary within each scene.
Co-training scenes and evaluation scenes share no scenes or objects (paper Section 5.1). Whether the DROID pretraining data contain similar scenes is not stated.
Scoring and access
- Scored by
- Progress score, Success rate113
More
Each task has a step rubric (reach, lift, place) normalised to 0-1. Simulation scores from object states; real rollouts were scored by humans with the same rubric. The code also logs binary success.
- Score
- Task progress score from 0 to 11220
More
Main score: normalised task progress (0-1) averaged over rollouts after co-training. Agreement with real robots is reported as Pearson r and MMRV.
An outside user asked whether the site's plots show progress or success (issue #8); the site states progress.
- Trials
- 50 per task in the paper1210+1
More
Paper · 50 simulated rollouts per task; 20 real rollouts per policy and scene for validation.1
RSS version · Bootstrap: 50 episodes resampled per policy, 500 repeats, 95% intervals.2
Public code · An outside user's run of the eval script covered 100 episodes per environment; the author says the UW scenes got 50 extra reset states for the release.10
CoVer-VLA · 50 episodes x 3 seeds on three PolaRiS environments.5
- Error bars
- Sometimes reported Inferred215
More
The RSS version adds bootstrap intervals for the correlation metrics; ablations show error bars over 5 seeds. Per-policy simulation scores are not given with intervals. CoVer-VLA reports standard deviations.
- Leaderboard
- None. Scores are only in papers. Inferred1513
More
The project site shows the paper's results; no submission page or ranking.
- Code licence
- MIT8
More
LICENSE file, copyright 2025 Arhan Jain.
- Data licence
- MIT for the PolaRiS Hub environments and the co-training dataset (Hugging Face cards)1912
- Asset licence
- MIT per the Hub card. Running PolaRiS needs Isaac Sim, whose Omniverse Kit SDK and NVIDIA 3D models are under NVIDIA's separate licence.192122
More
The Hub's robot folder (nvidia_droid: Franka arm and Robotiq gripper meshes) states no origin or separate terms; the folder name suggests NVIDIA (inferred).
- Access
- Open, on GitHub and Hugging Face131912+1
More
Code on GitHub; environments (<2 GB) and co-training data on ungated Hugging Face; co-trained checkpoints in Physical Intelligence's public Google Cloud bucket. Needs an NVIDIA GPU (tested on RTX 3090 and 5090) and Isaac Sim.
- Commercial use
- Unclear Inferred81921+1
More
Code and data are MIT. Isaac Sim's additional components are under NVIDIA's licence, which limits use to Isaac Sim and Isaac Lab, forbids redistribution and modification, is revocable and does not address commercial use explicitly (read 2026-10-10). The robot model's origin is unstated. Not legal advice.
Sources 23
- 1PolaRiS paper, full text v2 with figures (Figs. 6, 7, 8, 10, 13; Appendix C-D)Paper · Dec 2025 · checked 10 Oct 2026
- 2PolaRiS, RSS 2026 proceedings PDF (revised numbers, bootstrap intervals, Figs. 5-6, 10-11)Paper · Jul 2026 · checked 10 Oct 2026
- 3SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation (Section 5.1, Tables G.1-G.2, Appendix J)Paper · Jun 2026 · checked 10 Oct 2026
- 4PolaRiS issue #5: roboarena results values (owner reply)Repository · Jan 2026 · checked 10 Oct 2026
- 5Scaling Verification Can Be More Effective than Scaling Policy Learning for VLA Alignment (CoVer-VLA; Table 1 on PolaRiS)Paper · Feb 2026 · checked 10 Oct 2026
- 6H2RBench: A Real-to-Sim Benchmark for Evaluating Human-to-Robot Transfer (Section 3.3)Paper · Sep 2026 · checked 10 Oct 2026
- 7RLinf documentation: PolaRiS EvaluationOfficial site · 2026 · checked 10 Oct 2026
- 8PolaRiS LICENSE (MIT)Repository · Nov 2025 · checked 10 Oct 2026
- 9PolaRiS issue #13: Poor Env Init Poses (owner replies)Repository · Mar 2026 · checked 10 Oct 2026
- 10PolaRiS issue #24: Reproducing the results (pi0.5) (owner reply 2026-06-24)Repository · Jun 2026 · checked 10 Oct 2026
- 11Hugging Face Hub API: owhan/PolaRiS-Hub (downloads, likes, commits, path history)Index · 10 Oct 2026 · checked 10 Oct 2026
- 12PolaRiS co-training dataset card (license: mit) and dataset_info.jsonRepository · Dec 2025 · checked 10 Oct 2026
- 13PolaRiS README (environments, checkpoints, cotraining)Repository · Jul 2026 · checked 10 Oct 2026
- 14PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (arXiv abstract page; v1 2025-12-18, v2 2025-12-30)Paper · Dec 2025 · checked 10 Oct 2026
- 15PolaRiS project siteOfficial site · Dec 2025 · checked 10 Oct 2026
- 16PolaRiS, RSS 2026 proceedings pagePaper · Jul 2026 · checked 10 Oct 2026
- 17PolaRiS pyproject.toml (version 0.1.0; isaaclab[all,isaacsim]==2.3.0)Repository · Dec 2025 · checked 10 Oct 2026
- 18PolaRiS commit history, tags and releasesRepository · 13 Jul 2026 · checked 10 Oct 2026
- 19PolaRiS Hub dataset card (license: mit) and file treeRepository · Mar 2026 · checked 10 Oct 2026
- 20PolaRiS issue #8: Precise data in Figure 6Repository · Jan 2026 · checked 10 Oct 2026
- 21Isaac Sim LICENSE (Apache-2.0 plus NVIDIA additional components)Repository · 2025 · checked 10 Oct 2026
- 22NVIDIA Isaac Sim Additional Software and Materials LicenseOfficial site · 9 Jun 2025 · checked 10 Oct 2026
- 23PolaRiS docs: checkpoints_and_envs.mdRepository · Dec 2025 · checked 10 Oct 2026
Where we searched for missing information
validity (paired sim-vs-real studies of PolaRiS): PolaRiS arXiv v1, v2 and RSS 2026 PDF; SimFoundry 2606.28276; H2RBench 2609.24778 (built on PolaRiS, reports its own correlation); CoVer-VLA 2602.12281 (no real pairing on PolaRiS tasks); RoboWorld 2607.01060 (cites only); VLA Foundry 2604.19728 and IndustrialVLA-Bench 2609.25562 (no PolaRiS use found); PolaRiS GitHub issues; extended web searches.
PolaRiS comparison with SimplerEnv: PolaRiS v2 and RSS text: no SimplerEnv run; it states SIMPLER cannot evaluate wrist-camera policies and cites OpenVLA for SIMPLER's weak correlation, but OpenVLA v1-v3 do not mention SIMPLER (arXiv 2406.09246).
leaderboard: Project site, README, Hub card.
top_score: No single headline score; per-policy progress scores appear only in figures. Skipped.
Change history
- Created at full depth from primary sources (no basic entry existed). Includes the RSS 2026 revisions, the SimFoundry independent check and the public-environment reproduction problems from the GitHub issues.
- Published as a full entry.