VLN-CE

Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

How to read this picture

VLN-CE is a simulated benchmark in which an agent follows written route instructions through 3D scans of real buildings, using small, robot-like moves. It is widely used to test language-guided navigation models.123

Sources
Last checked 10 Oct 2026Full entry55 of 66 facts checked at the sourceNext check 8 Apr 2027
Runs in
Simulation2
Checked against real robots
Not checked
Skill
Navigation
Robot
Wheels only2
None (simulated LoCoBot-like agent)
Used by
Reported by most navigation models from 2025 to 2026456+2
683 citations
Licence
MIT9
Commercial use: not allowed

What a score here does not tell you Inferred

  1. How well an agent will do on a real robot.We found no study that compares VLN-CE scores with real-robot results for the same agents.
  2. How well an agent handles physical motion and collisions.The moves are idealised. In the default R2R setup, the agent slides along walls when it hits them.
  3. Whether a score can be fairly compared with scores in other papers.Sensors, training data and paper versions differ between papers.
ChartPublished scores over time
95% AND ABOVE203040506070809010020212022202320242025CMA_PM_DA_Aug (baseline) · 27.61% · 2020-12HPN+DN · 31.75% · 2021-03CWP-VLNBERT · 41.93% · 2021-10Reborn · 49.3% · 2022-04ETPNav · 55.13% · 2023-02BEVBert · 58.54% · 2023-05SRVLN · 63.91% · 2024-10NaVid · 45.1% · 2024-12VLN-CLASH · 66.4% · 2025-06ETP-R1 · 63.67% · 2025-08CoMaR (g3D-LF) · 62.32% · 2025-09CMA_PM_DA_Aug 27.61%VLN-CLASH 66.4%
Each dot is the average score reported in one paper. The shaded band marks the top 5% of the scale, where little room for improvement is left. A hollow dot means the model was trained with reinforcement learning inside the test environment.43

Comparisons with real robots

Tried on real robots, not compared Inferred101112+1

Details

VLN-PE (ICCV 2025) ran VLN-CE models with physically simulated humanoid, quadruped and wheeled robots in Isaac Sim and found a 34% relative drop in success; on a real Unitree Go2 over 14 episodes, a CMA model trained only on VLN-CE reached 7.14% success (28.57% after VLN-PE fine-tuning). NaVILA tested on a Unitree Go2 and a Booster T1 with 25 instructions repeated three times, comparing against GPT-4o, which it did not score on VLN-CE. NaVid also reports real-robot tests. Anderson et al. (CoRL 2020) ran a nav-graph R2R agent on a TurtleBot2 in an office with a scanned replica (46.8% real against 55.9% in simulation with a prepared map; 22.5% without), but that agent was trained on discrete R2R, not VLN-CE. See searched.

Our assessment Opinion

VLN-CE is closer to real robots than R2R. It has not been checked against real-robot results.

Reasoning

A VLN-CE score shows how well an agent follows route instructions through 3D scans of real buildings with small, robot-like moves. It is closer to a real robot than the original Room-to-Room (R2R) task, which moves between fixed viewpoints on a graph. Its motion is still idealised, and no study has paired VLN-CE scores with real-robot results.

Confidence: medium

Check the split, sensors and training data before comparing scores.

Reasoning

Before comparing two VLN-CE numbers, check the split, the camera setup (single RGB, panoramic, depth, odometry, waypoint predictor), the extra training data and the paper version. Published tables mix all of these.

Confidence: high

Read small gains on val-unseen with care.

Reasoning

Since January 2026, new results are self-reported on val-unseen, whose answers are public. Gains of one or two points should be read with care.

Confidence: medium

VLN-CE is widely used. Models trained on it are limited to academic use.

Reasoning

VLN-CE is the common benchmark for language-guided navigation models in 2025 and 2026, including company models from Alibaba. Its Matterport licence limits commercial use of models trained on it.

Confidence: medium

Known problems 5

  1. The same model appears with different scores

    Papers quote NaVid's success on the val-unseen split (buildings not seen in training) as 37.4% and 41.9%.12614+4

    Details

    NaVid reports 37.4% val-unseen success; Qwen-RobotNav's table lists NaVid at 41.9%. StreamVLN with extra data reported 56.9% in v1 (2025-07) and 56.4% in v2 (2026-07); later papers still cite 56.9%. ABot-N1 reports 70.9% (multi-task) and 68.3% (single-task); LightNav-0 lists ABot-N1 at 68.3%. The CMA baseline was published at 0.30 SPL on val-unseen but scores 0.27 on the leaderboard, which the README attributes to hardware and Habitat build differences.

  2. Methods with different sensors and data share one ranking

    Methods that use a single camera and methods that use a panoramic camera are ranked together.1157+2

    Details

    Some methods use a panoramic RGB-D camera, odometry and a waypoint predictor trained in the simulator; others use one forward RGB camera. Many recent models also train on extra VLN data beyond R2R-CE and RxR-CE (marked with a dagger in StreamVLN's table). The NaVILA, StreamVLN and ABot-N1 tables label these inputs, but the EvalAI leaderboard ranks all entries together. The RxR-Habitat rules required one 640 x 480 RGB-D camera, which excluded panoramic waypoint models.

  3. Results now come from a split with public answers

    The test server was retired in January 2026. New scores come from the val-unseen split, whose answers are public.1516

    Details

    In January 2026 the organisers retired the EvalAI test server and recommended reporting val-unseen, following RxR's test leaderboard. Ground-truth paths for val-unseen ship with the data (the test split's goals were hidden). The same split can therefore be used for model selection and for the reported score.

  4. The movement is idealised

    In VLN-PE, agents lost about a third of their success when they had to move as physically simulated robots.101718+1

    Details

    VLN-PE found that VLN-CE models lose 34% of their success, in relative terms, when they must move as physically simulated robots. The default R2R-CE configuration lets the agent slide along walls on collision (ALLOW_SLIDING: True), while the RxR-Habitat configuration disables sliding. In Habitat PointNav, Kadian et al. showed that sliding inflated simulated success; whether it does so in VLN-CE has not been measured.

  5. Early baselines did nearly as well without the instruction

    In the 2020 paper, a model given no instruction reached 17% success, against 20% for the full model.2

    Details

    In the original paper, a model given no instruction reached 17% val-unseen success and a model given no RGB image also reached 17%, against 20% for the full sequence-to-sequence baseline. The authors read this as shared path regularities between R2R and VLN-CE. Later models score far higher, and we found no newer ablation of this kind on VLN-CE.

Details

About

What it is
Benchmark12015
More

The paper and project site define a task with fixed splits and metrics; a public test server ran on EvalAI from January 2021.

Built by
Oregon State University, Georgia Tech, Facebook AI Research220
More

The RxR data come from Google Research. The RxR-Habitat Challenge was hosted by Oregon State University, Google Research and Meta AI (README).

Oregon State University · Jacob Krantz and Stefan Lee; the corresponding organiser of the challenges.23

Georgia Tech · Erik Wijmans, Arjun Majumdar, Dhruv Batra2

Facebook AI Research · Second affiliation of Erik Wijmans and Dhruv Batra2

Released
April 2020, at ECCV 2020120
More

arXiv v1 on 2020-04-06; code released in April 2020. Published at ECCV 2020.

The GitHub repository was created on 2020-04-03 (GitHub API).

Version
R2R_VLNCE v1-3 and RxR_VLNCE v0163
More

Current data: R2R_VLNCE_v1-3 (the README recommends this version) and RxR_VLNCE_v0. The code targets Habitat-Sim and Habitat-Lab 0.1.7 and Python 3.6.

R2R_VLNCE_v1-3 · Train 10,819 episodes (61 scenes); val seen 778 (53); val unseen 1,839 (11); test 3,408 (18). A preprocessed version adds 146,304 augmented episodes (envdrop).16

RxR_VLNCE_v0 · Train 60,300 episodes (59 scenes); val seen 6,746 (57); val unseen 11,006 (11); test-challenge 9,557 (17); roughly equal English, Hindi and Telugu.21322

Waypoint models (2021-10) · The repository added waypoint-based baselines that use panoramic observations; the README says these are not valid RxR-Habitat submissions.323

Last update
January 2026, when the test server was retired15244+1
More

In January 2026 the organisers announced that the EvalAI test server is being sunset; the challenge end date is 2026-01-31 and the newest test entry is dated 2026-01-17. The last code commit is 2025-01-07 (RxR data link fix).

Status
No code changes since January 2025. Use is very active. Inferred231525+1
More

Code unchanged since 2025-01-07 and the test server retired in January 2026. Use in new papers is very active.

29 issues are open (GitHub API). The RxR repository was archived on 2026-04-19 (s34). For current use see facts.used_by.

Setup

Runs in
Simulation2
Simulator
Habitat-Sim 0.1.732
More

Habitat-Sim and Habitat-Lab 0.1.7 with Matterport3D scene meshes

Robot
Wheels only2
More

The paper models 'a ground-based, zero-turning radius robot with a single, forward-mounted RGBD camera, similar to a LoCoBot'.

Robot model
No robot model. R2R setting: forward 0.25 m, turn 15 degrees, stop; RGB-D camera with a 90-degree field of view. RxR-Habitat setting: 30-degree turns, look up and down, 640 x 480 RGB-D.21718
More

The paper gives 256 x 256 RGB-D; the R2R config sets RGB to 224 x 224 and depth to 256 x 256.

Setting
Whole home Inferred23
More

We did not check the share of non-residential buildings in Matterport3D.

Tasks
16,844 R2R and 87,609 RxR episodes Inferred1621
More

16,844 R2R episodes (instruction plus path) over four splits. RxR_VLNCE adds 87,609 episodes.

Sums by us: 10,819 + 778 + 1,839 + 3,408 for R2R; 60,300 + 6,746 + 11,006 + 9,557 for RxR.

Scenes
90 scanned buildings2163
More

90 Matterport3D scenes. R2R splits: train 61, val seen 53, val unseen 11, test 18 scenes.

Training data
4,475 paths ported from R2R21627+1
More

4,475 R2R paths ported to continuous scenes (77% of R2R paths were navigable), each with about three instructions. Reference action sequences come with train and validation episodes; 146,304 augmented episodes are provided.

A ported path averages 55.88 low-level steps, against 4 to 6 hops in nav-graph R2R (paper Section 3.2). The paper found much lower scores than in nav-graph R2R: random agents reach about 3% success against 16.3% in R2R. Sim-2-Sim (ECCV 2022) transferred a nav-graph agent into VLN-CE and gained 12 points of success but did not keep its nav-graph performance.

Changes at test
New buildings and new instructions Inferred1621
More

Val-unseen (11 scenes) and test (18 scenes) use buildings not seen in training, with new instructions. RxR adds Hindi and Telugu instructions.

Scoring and access

Scored by
Success rate, Path efficiency152
More

RxR-Habitat ranks by nDTW, a path-similarity score with no taxonomy value.

Score
Success (stopping within 3 m of the goal) and SPL (success weighted by path length)152117
More

Success: the agent calls stop within 3 m of the goal, measured along walkable space. SPL: success weighted by path length; a successful episode scores the shortest-path length divided by the longer of the agent's path and the shortest path. Navigation error: distance left to the goal in metres. Oracle success: success if the agent had stopped at its closest point to the goal. RxR-Habitat ranks by nDTW, which scores from 0 to 1 how closely the agent's path follows the reference path.

EvalAI: the 3 m threshold is below the 5 m minimum start-to-goal distance in R2R.

Trials
One run per episode: 1,839 val-unseen or 3,408 test episodes. R2R episodes stop after 500 steps.1617
Who runs it
Both teams and organisers154
More

The agent runs on the team's machine; the server only scores the trajectory file.

Error bars
Not reported Inferred4116
More

The leaderboard and the papers we read report single numbers.

Leaderboard
Official, on EvalAI. It closed in January 2026.41524
More

EvalAI 'VLN-CE Challenge' test leaderboard: 46 public entries from 2020-12-28 to 2026-01-17, now closed. The RxR-Habitat leaderboard sits on a Google page that we could not read.

Code licence
MIT9
More

LICENSE file: MIT, copyright 2020 the five authors.

Data licence
CC BY-NC-SA 3.0 US, plus the Matterport3D terms32930
More

VLN-CE episode datasets and trained models: CC BY-NC-SA 3.0 US plus the Matterport3D Terms of Use. RxR instruction annotations: CC BY 4.0. Original R2R data: Matterport3D Terms of Use.

VLN-CE README: task datasets and trained models 'are considered data derived from the mp3d scene dataset'. The Matterport3D Simulator README puts R2R under the Matterport3D Terms of Use.

Asset licence
Matterport terms for academic use only3132
More

Matterport3D scenes: Matterport's End User License Agreement for Academic Use. Non-commercial academic use only; models trained on the data count as derived information and may not be used for non-academic purposes.

Access
By application Inferred332
More

Scenes: sign the Matterport3D terms and email them to receive a download script. Episode files and code are open downloads (Google Drive and GitHub).

Commercial use
Not allowed Inferred331
More

Episode data CC BY-NC-SA 3.0 US and scenes under Matterport's academic-only licence, which also covers trained models. Code MIT. Not legal advice.

Sources 32

  1. 1Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (arXiv abstract page)Paper · Apr 2020 · checked 10 Oct 2026
  2. 2VLN-CE paper, full text v2 (Sections 3 to 5, Tables 2 to 4)Paper · May 2020 · checked 10 Oct 2026
  3. 3VLN-CE GitHub README (data, RxR-Habitat Challenge, baseline performance, licence)Repository · Jan 2025 · checked 10 Oct 2026
  4. 4EvalAI: VLN-CE Challenge test leaderboard (phase split 1966; 46 public entries)Leaderboard · 17 Jan 2026 · checked 10 Oct 2026
  5. 5StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling, v2 (Table I)Paper · Jul 2025 · checked 10 Oct 2026
  6. 6Qwen-RobotNav Technical Report (Table 1: VLN-CE val-unseen)Paper · Jun 2026 · checked 10 Oct 2026
  7. 7ABot-N1: Toward a General Visual Language Navigation Foundation Model (Table 1)Paper · Jul 2026 · checked 10 Oct 2026
  8. 8LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation (Table III)Paper · Aug 2026 · checked 10 Oct 2026
  9. 9VLN-CE LICENSE fileRepository · 2020 · checked 10 Oct 2026
  10. 10Rethinking the Embodied Gap in Vision-and-Language Navigation (VLN-PE), v2Paper · Jul 2025 · checked 10 Oct 2026
  11. 11NaVILA: Legged Robot Vision-Language-Action Model for Navigation (Tables I and VI)Paper · Dec 2024 · checked 10 Oct 2026
  12. 12NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation (v7, Table I)Paper · Feb 2024 · checked 10 Oct 2026
  13. 13Sim-to-Real Transfer for Vision-and-Language Navigation (Anderson et al., CoRL 2020)Paper · Nov 2020 · checked 10 Oct 2026
  14. 14StreamVLN, v1 (Table I)Paper · Jul 2025 · checked 10 Oct 2026
  15. 15EvalAI: VLN-CE Challenge (challenge 719) description, evaluation and submission guidelines, with the January 2026 sunset noticeLeaderboard · Jan 2026 · checked 10 Oct 2026
  16. 16VLN-CE dataset page (R2R_VLNCE_v1-3 split counts and format)Official site · 2021 · checked 10 Oct 2026
  17. 17VLN-CE R2R task configuration (vlnce_task.yaml: ALLOW_SLIDING True, 500 steps, 3.0 m success)Repository · 2021 · checked 10 Oct 2026
  18. 18VLN-CE RxR English task configuration (ALLOW_SLIDING False, 30-degree turns, 640 x 480 RGB-D)Repository · 2021 · checked 10 Oct 2026
  19. 19Sim2Real Predictivity (Kadian et al.): wall sliding in Habitat inflated simulated PointNav successPaper · Dec 2019 · checked 10 Oct 2026
  20. 20VLN-CE project site (news, people, leaderboard link)Official site · Jan 2021 · checked 10 Oct 2026
  21. 21Retrospectives on the Embodied AI Workshop (RxR-Habitat section)Paper · Oct 2022 · checked 10 Oct 2026
  22. 22Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingPaper · Oct 2020 · checked 10 Oct 2026
  23. 23VLN-CE commit historyRepository · 7 Jan 2025 · checked 10 Oct 2026
  24. 24EvalAI API: VLN-CE Challenge phases (test: 1 per day, 5 in total; end 2026-01-31)Leaderboard · 31 Jan 2026 · checked 10 Oct 2026
  25. 25GitHub API: jacobkrantz/VLN-CE (stars, forks, created, pushed)Index · 10 Oct 2026 · checked 10 Oct 2026
  26. 26GitHub GraphQL: google-research-datasets/RxR archivedAt 2026-04-19Index · 10 Oct 2026 · checked 10 Oct 2026
  27. 27Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments (R2R)Paper · Nov 2017 · checked 10 Oct 2026
  28. 28Sim-2-Sim Transfer for Vision-and-Language Navigation in Continuous EnvironmentsPaper · Apr 2022 · checked 10 Oct 2026
  29. 29Room-Across-Room (RxR) repository README and LICENSE (CC BY 4.0; archived 2026-04-19 per GitHub)Repository · Jul 2023 · checked 10 Oct 2026
  30. 30Matterport3D Simulator README (licence: Matterport3D data and derived data under the Matterport3D Terms of Use; code MIT)Repository · Jul 2024 · checked 10 Oct 2026
  31. 31Matterport End User License Agreement for Academic Use of Model DataOfficial site · unknown · checked 10 Oct 2026
  32. 32Matterport3D Terms of Use (PDF linked from the VLN-CE README)Official site · unknown · checked 10 Oct 2026
Where we searched for missing information

sim_to_real: VLN-CE paper and project site; Sim-2-Sim (2204.09667); Anderson et al. CoRL 2020 (discrete R2R agent); VLN-PE (2507.13019, Isaac Sim plus 14 real episodes); NaVid, NaVILA, StreamVLN real-robot sections; web searches on 2026-10-10 for VLN-CE sim-to-real correlation and real-robot evaluation. No study pairs VLN-CE scores and real-robot results across several policies.

RxR-Habitat leaderboard: ai.google.com/research/rxr/habitat (JavaScript page; no readable content via curl or fetch). Results taken from the organisers' retrospective, the 2022 winner report and the CVPR 2023 workshop page.

license_data (R2R annotations): bringmeaspoon.org (no licence statement); Matterport3D Simulator README (data derived from Matterport3D under its Terms of Use).

citations: Semantic Scholar API, repeated HTTP 429 responses on 2026-10-10.

Change history

  1. Created as a full entry from primary sources. Covers R2R-CE and RxR-CE (RxR-Habitat). Added the January 2026 test-server sunset, the official leaderboard history (46 public entries), licences (CC BY-NC-SA 3.0 US data, Matterport academic scene licence, CC BY 4.0 RxR annotations) and five issues.