RoboWorld Challenge 2026 · Track #3

SafeDrive-VLA

Towards Safety in Autonomous Driving

Language-guided driving that follows safe instructions and responds safely to instruction-scene conflicts.

Benchmark
CARLA-F / B2D-C
Setting
Closed-loop safety-aware driving
Metrics
NCR, speed error; DS, SR, infractions

Introduction

Participants develop vision-language-action models (VLAs) for safe navigation guided by natural-language instructions. Given onboard visual observations and an instruction, models generate driving trajectories or control actions that account for surrounding traffic conditions.

Models should follow safe instructions, including turns, lane changes, and target-speed requests. When an instruction conflicts with the traffic scene, models must respond appropriately by slowing down, waiting, or selecting a safe alternative.

Reference

SafeDriveVLA: Navigation-Conditioned World Model Dreaming for Conflict-Aware End-to-End Autonomous Driving

CoRL 2026

S. Xie, Z. Zhang, J. Wang, J. Qu, X. Liang, L. Kong, J. Lu, H. I. Christensen, Q. A. Chen

Registration

Register Your Team  (Google Form)

Submission

Submission Portal  (CodaBench Server)

Dataset & Evaluation Setting

Evaluation consists of 40 closed-loop simulation routes across two complementary test sets.

Split / Test SetSizeDetails
CARLA-F20 routesSafe instruction following
B2D-C20 routesConflicting instructions issued when hazards occur

Training Data

Training uses the CARLA driving dataset that SimLingo (CVPR 2025) collected with PDM-Lite, a privileged rule-based expert. The dataset is released on Hugging Face as RenzKa/simlingo.

Evaluation

CARLA-F evaluates safe instruction following using per-instruction Navigation Compliance Rate (NCR) and target-speed error in m/s.

B2D-C evaluates safety-critical driving using Driving Score (DS), Success Rate (SR), and counts of collisions, traffic violations, and route departures. Results from the two test sets will be reported separately.

MetricPreferred DirectionDefinition
CARLA-F: Navigation Compliance Rate (NCR)HigherPer-instruction navigation compliance.
CARLA-F: target-speed error (m/s)LowerDeviation from the instructed target speed.
B2D-C: Driving Score (DS) and Success Rate (SR)HigherClosed-loop driving performance and route success.
B2D-C: collisions, traffic violations, route departuresLowerCounts of safety and route-following failures.

Reference Baselines

SimLingo and SafeDriveVLA both generate driving actions from visual observations and language instructions. SafeDriveVLA additionally uses explicit driving-mode selection and world-model predictions to handle conflicts between instructions and the traffic scene.

SafeDriveVLA overview. Left: world-model pre-training, where a frozen V-JEPA 2 encoder embeds camera frames and a world state predictor forecasts future latents from action and state tokens. Right: the SafeDriveVLA policy, where InternVL3-1B reads interleaved text, visual, and world tokens and emits a mode token, a begin-of-action token, and action tokens.

SafeDriveVLA decouples conflict reasoning from action generation:

The SafeDriveVLA repository provides world-model pre-training, VLA training, and closed-loop evaluation code for CARLA-F and B2D-C, together with the benchmark route files and data-preparation instructions.

Reference benchmark results are shown below; these are baseline results reported in the SafeDriveVLA paper on the full benchmark routes, not competition submissions.

BaselineCARLA-F (NCR)B2D-C (DS)
SimLingo52.4%72.8
SafeDriveVLA82.7%67.3

The CARLA-F values are the average Navigation Compliance Rate (NCR), and the B2D-C values are the average Driving Score (DS). The two test sets measure different aspects of performance: (1) instruction following in safe scenarios and (2) awareness of unsafe instructions.

Track Organizer

Contact

For challenge questions, contact roboworld2026@gmail.com.