RoboWorld Challenge 2026 · Track #1
WorldNav
Language-Conditioned World Navigation
Open-loop visual navigation from one initial RGB observation and a natural-language instruction.
- Benchmark
- LCVN
- Setting
- Open-loop navigation
- Metrics
- SR, ATE, RPE, Full-route fidelity, Heading and turn fidelity
Introduction
Participants develop world-model-based or vision-language-action (VLA) agents for language-conditioned visual navigation. Given a single initial egocentric RGB observation and a natural-language instruction, models generate the full future trajectory as continuous actions: planar displacement and yaw rotation, terminated by a Stop command or a null action.
No goal image is provided, and no intermediate environmental feedback is available. Agents must ground the instruction in the initial observation and plan in an open-loop manner.
The track encourages methods that couple imagination with control, including models that predict future observations to select actions, autoregressive models that interleave observation and action prediction, and VLAs that map vision and language directly to continuous actions. Policy-only baselines are also permitted.
Reference
Language-Conditioned World Modeling for Visual Navigation
NeurIPS 2026 Oral
Registration
Register Your Team (Google Form)Submission
Submission Portal (CodaBench Server)Dataset & Evaluation Setting
LCVN contains 39,016 trajectories and 117,048 human-verified instructions sourced from Go Stanford, ReCon, SCAND, HuRoN, and TartanDrive. Each trajectory has three instruction styles: concise, intricate, and landmark-grounded.
| Split / Test Set | Size | Details |
|---|---|---|
| Training | 28,813 trajectories | 86,439 instructions |
| Validation: seen | 3,602 trajectories | Seen validation environments |
| Validation: unseen | 1,500 trajectories | Drawn exclusively from TartanDrive |
| Held-out test | 5,101 trajectories | TartanDrive and in-domain sources; annotations withheld |
Evaluation
All submissions are scored on the held-out test split. Navigation quality is assessed using Success Rate (SR), Absolute Trajectory Error (ATE), Relative Pose Error (RPE), Full-route fidelity, and Heading and turn fidelity.
A trajectory is successful when its final distance to the goal is below the agent's average step size. Submissions are ranked by the composite Score defined below, with higher scores indicating better performance.
| Metric | Preferred Direction | Definition |
|---|---|---|
| Success Rate (SR) | Higher | Fraction of trajectories satisfying the goal-distance success criterion. |
| Absolute Trajectory Error (ATE) | Lower | Absolute trajectory error against the reference trajectory. |
| Relative Pose Error (RPE) | Lower | Relative pose error against the reference trajectory. |
| Full-route fidelity | Higher | Similarity of the entire ordered predicted route to the recorded route, with excess travel penalized. |
| Heading and turn fidelity | Higher | Agreement of facing direction during travel and of the direction, location, and amount of turns. |
Composite Score
The final ranking uses the following Score, averaged over all evaluated episodes:
Here, N is the number of evaluated episodes. Fi and Hi denote Full-route fidelity and Heading and turn fidelity for episode i, respectively. SRi is binary success; ATEi and RPEi are per-episode errors in dataset coordinate units.
Reference Baselines
LCVN-Uni is a unified autoregressive multimodal model fine-tuned from Anole-7B. It jointly predicts the next action and next observation in a single forward pass.
LCVN-WM + LCVN-AC pairs a language-conditioned diffusion world model with an actor-critic agent trained entirely in the world model's latent space.
Reference benchmark results are shown below; these are baseline results, not competition submissions.
| Baseline | SR | ATE | RPE |
|---|---|---|---|
| LCVN-Uni | 0.36 | 0.65 | 0.22 |
| LCVN-WM + LCVN-AC | 0.34 | 0.76 | 0.26 |
Track Organizer
Contact
For challenge questions, contact roboworld2026@gmail.com.
