RoboWorld Challenge 2026 · Track #1

WorldNav

Language-Conditioned World Navigation

Open-loop visual navigation from one initial RGB observation and a natural-language instruction.

Benchmark
LCVN
Setting
Open-loop navigation
Metrics
SR, ATE, RPE, Full-route fidelity, Heading and turn fidelity

Introduction

Participants develop world-model-based or vision-language-action (VLA) agents for language-conditioned visual navigation. Given a single initial egocentric RGB observation and a natural-language instruction, models generate the full future trajectory as continuous actions: planar displacement and yaw rotation, terminated by a Stop command or a null action.

No goal image is provided, and no intermediate environmental feedback is available. Agents must ground the instruction in the initial observation and plan in an open-loop manner.

The track encourages methods that couple imagination with control, including models that predict future observations to select actions, autoregressive models that interleave observation and action prediction, and VLAs that map vision and language directly to continuous actions. Policy-only baselines are also permitted.

Reference

Y. Dong, F. Wu, Y. Dai, L. Kong, G. Chen, Y. Sha, Q. Hu, F. Liu, S. Huang, Q. Dai, Z.-Q. Cheng

Registration

Register Your Team  (Google Form)

Submission

Submission Portal  (CodaBench Server)

Dataset & Evaluation Setting

LCVN contains 39,016 trajectories and 117,048 human-verified instructions sourced from Go Stanford, ReCon, SCAND, HuRoN, and TartanDrive. Each trajectory has three instruction styles: concise, intricate, and landmark-grounded.

Split / Test SetSizeDetails
Training28,813 trajectories86,439 instructions
Validation: seen3,602 trajectoriesSeen validation environments
Validation: unseen1,500 trajectoriesDrawn exclusively from TartanDrive
Held-out test5,101 trajectoriesTartanDrive and in-domain sources; annotations withheld

Evaluation

All submissions are scored on the held-out test split. Navigation quality is assessed using Success Rate (SR), Absolute Trajectory Error (ATE), Relative Pose Error (RPE), Full-route fidelity, and Heading and turn fidelity.

A trajectory is successful when its final distance to the goal is below the agent's average step size. Submissions are ranked by the composite Score defined below, with higher scores indicating better performance.

MetricPreferred DirectionDefinition
Success Rate (SR)HigherFraction of trajectories satisfying the goal-distance success criterion.
Absolute Trajectory Error (ATE)LowerAbsolute trajectory error against the reference trajectory.
Relative Pose Error (RPE)LowerRelative pose error against the reference trajectory.
Full-route fidelityHigherSimilarity of the entire ordered predicted route to the recorded route, with excess travel penalized.
Heading and turn fidelityHigherAgreement of facing direction during travel and of the direction, location, and amount of turns.

Composite Score

The final ranking uses the following Score, averaged over all evaluated episodes:

Score = 100 N ∑ i=1 N Fi ( 1 + Hi ) 2 × ( 0.70 + 0.10 SRi + 0.10 1 + ATEi + 0.10 1 + RPEi )

Here, N is the number of evaluated episodes. Fi and Hi denote Full-route fidelity and Heading and turn fidelity for episode i, respectively. SRi is binary success; ATEi and RPEi are per-episode errors in dataset coordinate units.

Reference Baselines

LCVN-Uni is a unified autoregressive multimodal model fine-tuned from Anole-7B. It jointly predicts the next action and next observation in a single forward pass.

LCVN-WM + LCVN-AC pairs a language-conditioned diffusion world model with an actor-critic agent trained entirely in the world model's latent space.

Reference benchmark results are shown below; these are baseline results, not competition submissions.

BaselineSRATERPE
LCVN-Uni0.360.650.22
LCVN-WM + LCVN-AC0.340.760.26

Track Organizer

Contact

For challenge questions, contact roboworld2026@gmail.com.