VLN-CE
Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments
VLN-CE is a simulated benchmark in which an agent follows written route instructions through 3D scans of real buildings, using small, robot-like moves. It is widely used to test language-guided navigation models.123
What a score here does not tell you Inferred
- How well an agent will do on a real robot.We found no study that compares VLN-CE scores with real-robot results for the same agents.
- How well an agent handles physical motion and collisions.The moves are idealised. In the default R2R setup, the agent slides along walls when it hits them.
- Whether a score can be fairly compared with scores in other papers.Sensors, training data and paper versions differ between papers.
Comparisons with real robots
Tried on real robots, not compared Inferred101112+1
Details
VLN-PE (ICCV 2025) ran VLN-CE models with physically simulated humanoid, quadruped and wheeled robots in Isaac Sim and found a 34% relative drop in success; on a real Unitree Go2 over 14 episodes, a CMA model trained only on VLN-CE reached 7.14% success (28.57% after VLN-PE fine-tuning). NaVILA tested on a Unitree Go2 and a Booster T1 with 25 instructions repeated three times, comparing against GPT-4o, which it did not score on VLN-CE. NaVid also reports real-robot tests. Anderson et al. (CoRL 2020) ran a nav-graph R2R agent on a TurtleBot2 in an office with a scanned replica (46.8% real against 55.9% in simulation with a prepared map; 22.5% without), but that agent was trained on discrete R2R, not VLN-CE. See searched.
Our assessment Opinion
VLN-CE is closer to real robots than R2R. It has not been checked against real-robot results.
Reasoning
A VLN-CE score shows how well an agent follows route instructions through 3D scans of real buildings with small, robot-like moves. It is closer to a real robot than the original Room-to-Room (R2R) task, which moves between fixed viewpoints on a graph. Its motion is still idealised, and no study has paired VLN-CE scores with real-robot results.
Confidence: medium
Check the split, sensors and training data before comparing scores.
Reasoning
Before comparing two VLN-CE numbers, check the split, the camera setup (single RGB, panoramic, depth, odometry, waypoint predictor), the extra training data and the paper version. Published tables mix all of these.
Confidence: high
Read small gains on val-unseen with care.
Reasoning
Since January 2026, new results are self-reported on val-unseen, whose answers are public. Gains of one or two points should be read with care.
Confidence: medium
VLN-CE is widely used. Models trained on it are limited to academic use.
Reasoning
VLN-CE is the common benchmark for language-guided navigation models in 2025 and 2026, including company models from Alibaba. Its Matterport licence limits commercial use of models trained on it.
Confidence: medium
Known problems 5
The same model appears with different scores
Papers quote NaVid's success on the val-unseen split (buildings not seen in training) as 37.4% and 41.9%.12614+4
Details
NaVid reports 37.4% val-unseen success; Qwen-RobotNav's table lists NaVid at 41.9%. StreamVLN with extra data reported 56.9% in v1 (2025-07) and 56.4% in v2 (2026-07); later papers still cite 56.9%. ABot-N1 reports 70.9% (multi-task) and 68.3% (single-task); LightNav-0 lists ABot-N1 at 68.3%. The CMA baseline was published at 0.30 SPL on val-unseen but scores 0.27 on the leaderboard, which the README attributes to hardware and Habitat build differences.
Methods with different sensors and data share one ranking
Methods that use a single camera and methods that use a panoramic camera are ranked together.1157+2
Details
Some methods use a panoramic RGB-D camera, odometry and a waypoint predictor trained in the simulator; others use one forward RGB camera. Many recent models also train on extra VLN data beyond R2R-CE and RxR-CE (marked with a dagger in StreamVLN's table). The NaVILA, StreamVLN and ABot-N1 tables label these inputs, but the EvalAI leaderboard ranks all entries together. The RxR-Habitat rules required one 640 x 480 RGB-D camera, which excluded panoramic waypoint models.
Results now come from a split with public answers
The test server was retired in January 2026. New scores come from the val-unseen split, whose answers are public.1516
Details
In January 2026 the organisers retired the EvalAI test server and recommended reporting val-unseen, following RxR's test leaderboard. Ground-truth paths for val-unseen ship with the data (the test split's goals were hidden). The same split can therefore be used for model selection and for the reported score.
The movement is idealised
In VLN-PE, agents lost about a third of their success when they had to move as physically simulated robots.101718+1
Details
VLN-PE found that VLN-CE models lose 34% of their success, in relative terms, when they must move as physically simulated robots. The default R2R-CE configuration lets the agent slide along walls on collision (ALLOW_SLIDING: True), while the RxR-Habitat configuration disables sliding. In Habitat PointNav, Kadian et al. showed that sliding inflated simulated success; whether it does so in VLN-CE has not been measured.
Early baselines did nearly as well without the instruction
In the 2020 paper, a model given no instruction reached 17% success, against 20% for the full model.2
Details
In the original paper, a model given no instruction reached 17% val-unseen success and a model given no RGB image also reached 17%, against 20% for the full sequence-to-sequence baseline. The authors read this as shared path regularities between R2R and VLN-CE. Later models score far higher, and we found no newer ablation of this kind on VLN-CE.
Details
About
- What it is
- Benchmark12015
More
The paper and project site define a task with fixed splits and metrics; a public test server ran on EvalAI from January 2021.
- Built by
- Oregon State University, Georgia Tech, Facebook AI Research220
More
The RxR data come from Google Research. The RxR-Habitat Challenge was hosted by Oregon State University, Google Research and Meta AI (README).
Oregon State University · Jacob Krantz and Stefan Lee; the corresponding organiser of the challenges.23
Georgia Tech · Erik Wijmans, Arjun Majumdar, Dhruv Batra2
Facebook AI Research · Second affiliation of Erik Wijmans and Dhruv Batra2
- Released
- April 2020, at ECCV 2020120
More
arXiv v1 on 2020-04-06; code released in April 2020. Published at ECCV 2020.
The GitHub repository was created on 2020-04-03 (GitHub API).
- Version
- R2R_VLNCE v1-3 and RxR_VLNCE v0163
More
Current data: R2R_VLNCE_v1-3 (the README recommends this version) and RxR_VLNCE_v0. The code targets Habitat-Sim and Habitat-Lab 0.1.7 and Python 3.6.
R2R_VLNCE_v1-3 · Train 10,819 episodes (61 scenes); val seen 778 (53); val unseen 1,839 (11); test 3,408 (18). A preprocessed version adds 146,304 augmented episodes (envdrop).16
RxR_VLNCE_v0 · Train 60,300 episodes (59 scenes); val seen 6,746 (57); val unseen 11,006 (11); test-challenge 9,557 (17); roughly equal English, Hindi and Telugu.21322
Waypoint models (2021-10) · The repository added waypoint-based baselines that use panoramic observations; the README says these are not valid RxR-Habitat submissions.323
- Last update
- January 2026, when the test server was retired15244+1
More
In January 2026 the organisers announced that the EvalAI test server is being sunset; the challenge end date is 2026-01-31 and the newest test entry is dated 2026-01-17. The last code commit is 2025-01-07 (RxR data link fix).
- Status
- No code changes since January 2025. Use is very active. Inferred231525+1
More
Code unchanged since 2025-01-07 and the test server retired in January 2026. Use in new papers is very active.
29 issues are open (GitHub API). The RxR repository was archived on 2026-04-19 (s34). For current use see facts.used_by.
Setup
- Runs in
- Simulation2
- Robot
- Wheels only2
More
The paper models 'a ground-based, zero-turning radius robot with a single, forward-mounted RGBD camera, similar to a LoCoBot'.
- Robot model
- No robot model. R2R setting: forward 0.25 m, turn 15 degrees, stop; RGB-D camera with a 90-degree field of view. RxR-Habitat setting: 30-degree turns, look up and down, 640 x 480 RGB-D.21718
More
The paper gives 256 x 256 RGB-D; the R2R config sets RGB to 224 x 224 and depth to 256 x 256.
- Setting
- Whole home Inferred23
More
We did not check the share of non-residential buildings in Matterport3D.
- Tasks
- 16,844 R2R and 87,609 RxR episodes Inferred1621
More
16,844 R2R episodes (instruction plus path) over four splits. RxR_VLNCE adds 87,609 episodes.
Sums by us: 10,819 + 778 + 1,839 + 3,408 for R2R; 60,300 + 6,746 + 11,006 + 9,557 for RxR.
- Scenes
- 90 scanned buildings2163
More
90 Matterport3D scenes. R2R splits: train 61, val seen 53, val unseen 11, test 18 scenes.
- Training data
- 4,475 paths ported from R2R21627+1
More
4,475 R2R paths ported to continuous scenes (77% of R2R paths were navigable), each with about three instructions. Reference action sequences come with train and validation episodes; 146,304 augmented episodes are provided.
A ported path averages 55.88 low-level steps, against 4 to 6 hops in nav-graph R2R (paper Section 3.2). The paper found much lower scores than in nav-graph R2R: random agents reach about 3% success against 16.3% in R2R. Sim-2-Sim (ECCV 2022) transferred a nav-graph agent into VLN-CE and gained 12 points of success but did not keep its nav-graph performance.
Scoring and access
- Scored by
- Success rate, Path efficiency152
More
RxR-Habitat ranks by nDTW, a path-similarity score with no taxonomy value.
- Score
- Success (stopping within 3 m of the goal) and SPL (success weighted by path length)152117
More
Success: the agent calls stop within 3 m of the goal, measured along walkable space. SPL: success weighted by path length; a successful episode scores the shortest-path length divided by the longer of the agent's path and the shortest path. Navigation error: distance left to the goal in metres. Oracle success: success if the agent had stopped at its closest point to the goal. RxR-Habitat ranks by nDTW, which scores from 0 to 1 how closely the agent's path follows the reference path.
EvalAI: the 3 m threshold is below the 5 m minimum start-to-goal distance in R2R.
- Trials
- One run per episode: 1,839 val-unseen or 3,408 test episodes. R2R episodes stop after 500 steps.1617
- Who runs it
- Both teams and organisers154
More
The agent runs on the team's machine; the server only scores the trajectory file.
- Error bars
- Not reported Inferred4116
More
The leaderboard and the papers we read report single numbers.
- Leaderboard
- Official, on EvalAI. It closed in January 2026.41524
More
EvalAI 'VLN-CE Challenge' test leaderboard: 46 public entries from 2020-12-28 to 2026-01-17, now closed. The RxR-Habitat leaderboard sits on a Google page that we could not read.
- Code licence
- MIT9
More
LICENSE file: MIT, copyright 2020 the five authors.
- Data licence
- CC BY-NC-SA 3.0 US, plus the Matterport3D terms32930
More
VLN-CE episode datasets and trained models: CC BY-NC-SA 3.0 US plus the Matterport3D Terms of Use. RxR instruction annotations: CC BY 4.0. Original R2R data: Matterport3D Terms of Use.
VLN-CE README: task datasets and trained models 'are considered data derived from the mp3d scene dataset'. The Matterport3D Simulator README puts R2R under the Matterport3D Terms of Use.
- Asset licence
- Matterport terms for academic use only3132
More
Matterport3D scenes: Matterport's End User License Agreement for Academic Use. Non-commercial academic use only; models trained on the data count as derived information and may not be used for non-academic purposes.
Sources 32
- 1Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (arXiv abstract page)Paper · Apr 2020 · checked 10 Oct 2026
- 2VLN-CE paper, full text v2 (Sections 3 to 5, Tables 2 to 4)Paper · May 2020 · checked 10 Oct 2026
- 3VLN-CE GitHub README (data, RxR-Habitat Challenge, baseline performance, licence)Repository · Jan 2025 · checked 10 Oct 2026
- 4EvalAI: VLN-CE Challenge test leaderboard (phase split 1966; 46 public entries)Leaderboard · 17 Jan 2026 · checked 10 Oct 2026
- 5StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling, v2 (Table I)Paper · Jul 2025 · checked 10 Oct 2026
- 6Qwen-RobotNav Technical Report (Table 1: VLN-CE val-unseen)Paper · Jun 2026 · checked 10 Oct 2026
- 7ABot-N1: Toward a General Visual Language Navigation Foundation Model (Table 1)Paper · Jul 2026 · checked 10 Oct 2026
- 8LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation (Table III)Paper · Aug 2026 · checked 10 Oct 2026
- 9VLN-CE LICENSE fileRepository · 2020 · checked 10 Oct 2026
- 10Rethinking the Embodied Gap in Vision-and-Language Navigation (VLN-PE), v2Paper · Jul 2025 · checked 10 Oct 2026
- 11NaVILA: Legged Robot Vision-Language-Action Model for Navigation (Tables I and VI)Paper · Dec 2024 · checked 10 Oct 2026
- 12NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation (v7, Table I)Paper · Feb 2024 · checked 10 Oct 2026
- 13Sim-to-Real Transfer for Vision-and-Language Navigation (Anderson et al., CoRL 2020)Paper · Nov 2020 · checked 10 Oct 2026
- 14StreamVLN, v1 (Table I)Paper · Jul 2025 · checked 10 Oct 2026
- 15EvalAI: VLN-CE Challenge (challenge 719) description, evaluation and submission guidelines, with the January 2026 sunset noticeLeaderboard · Jan 2026 · checked 10 Oct 2026
- 16VLN-CE dataset page (R2R_VLNCE_v1-3 split counts and format)Official site · 2021 · checked 10 Oct 2026
- 17VLN-CE R2R task configuration (vlnce_task.yaml: ALLOW_SLIDING True, 500 steps, 3.0 m success)Repository · 2021 · checked 10 Oct 2026
- 18VLN-CE RxR English task configuration (ALLOW_SLIDING False, 30-degree turns, 640 x 480 RGB-D)Repository · 2021 · checked 10 Oct 2026
- 19Sim2Real Predictivity (Kadian et al.): wall sliding in Habitat inflated simulated PointNav successPaper · Dec 2019 · checked 10 Oct 2026
- 20VLN-CE project site (news, people, leaderboard link)Official site · Jan 2021 · checked 10 Oct 2026
- 21Retrospectives on the Embodied AI Workshop (RxR-Habitat section)Paper · Oct 2022 · checked 10 Oct 2026
- 22Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingPaper · Oct 2020 · checked 10 Oct 2026
- 23VLN-CE commit historyRepository · 7 Jan 2025 · checked 10 Oct 2026
- 24EvalAI API: VLN-CE Challenge phases (test: 1 per day, 5 in total; end 2026-01-31)Leaderboard · 31 Jan 2026 · checked 10 Oct 2026
- 25GitHub API: jacobkrantz/VLN-CE (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
- 26GitHub GraphQL: google-research-datasets/RxR archivedAt 2026-04-19Index · 10 Oct 2026 · checked 10 Oct 2026
- 27Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments (R2R)Paper · Nov 2017 · checked 10 Oct 2026
- 28Sim-2-Sim Transfer for Vision-and-Language Navigation in Continuous EnvironmentsPaper · Apr 2022 · checked 10 Oct 2026
- 29Room-Across-Room (RxR) repository README and LICENSE (CC BY 4.0; archived 2026-04-19 per GitHub)Repository · Jul 2023 · checked 10 Oct 2026
- 30Matterport3D Simulator README (licence: Matterport3D data and derived data under the Matterport3D Terms of Use; code MIT)Repository · Jul 2024 · checked 10 Oct 2026
- 31Matterport End User License Agreement for Academic Use of Model DataOfficial site · unknown · checked 10 Oct 2026
- 32Matterport3D Terms of Use (PDF linked from the VLN-CE README)Official site · unknown · checked 10 Oct 2026
Where we searched for missing information
sim_to_real: VLN-CE paper and project site; Sim-2-Sim (2204.09667); Anderson et al. CoRL 2020 (discrete R2R agent); VLN-PE (2507.13019, Isaac Sim plus 14 real episodes); NaVid, NaVILA, StreamVLN real-robot sections; web searches on 2026-10-10 for VLN-CE sim-to-real correlation and real-robot evaluation. No study pairs VLN-CE scores and real-robot results across several policies.
RxR-Habitat leaderboard: ai.google.com/research/rxr/habitat (JavaScript page; no readable content via curl or fetch). Results taken from the organisers' retrospective, the 2022 winner report and the CVPR 2023 workshop page.
license_data (R2R annotations): bringmeaspoon.org (no licence statement); Matterport3D Simulator README (data derived from Matterport3D under its Terms of Use).
citations: Semantic Scholar API, repeated HTTP 429 responses on 2026-10-10.
Change history
- Created as a full entry from primary sources. Covers R2R-CE and RxR-CE (RxR-Habitat). Added the January 2026 test-server sunset, the official leaderboard history (46 public entries), licences (CC BY-NC-SA 3.0 US data, Matterport academic scene licence, CC BY 4.0 RxR annotations) and five issues.