RoboWorld Challenge 2026 · Track #2

HA-VLN 2.0

Human-Aware Social Navigation

Closed-loop RGB-D navigation with socially grounded language and dynamic, interacting humans.

Benchmark
HA-R2R / HA-VLN 2.0
Setting
Closed-loop social navigation
Metrics
NE, SR, TCR, CR

Introduction

Participants develop agents for vision-and-language navigation in continuous, photorealistic 3D environments populated by dynamic, interacting humans. Given egocentric RGB-D observations and a natural-language instruction referencing ongoing human activities, agents must reach the described goal while maintaining socially compliant distances and avoiding collisions with moving bystanders.

The HA-VLN 2.0 simulator uses fine-grained actions: 0.25 m forward steps and 15-degree turns. Partial observability, interactions among multiple people, and unpredictable human motion require real-time re-planning when corridors or doorways become blocked.

Policy-based agents, imitation or reinforcement learning, waypoint planners, and topological planners are welcome. The track also encourages world models that anticipate human trajectories and future observations, and VLAs that ground human-centric instructions in egocentric video and output low-level actions.

Reference

Y. Dong, F. Wu, Q. He, L. Kong, H. Li, M. Li, Z. Cheng, Y. Zhou, J. Sun, Q. Dai, A. G. Hauptmann, Z.-Q. Cheng

Registration

Register Your Team  (Google Form)

Submission

Submission Portal  (CodaBench Server)

Dataset & Evaluation Setting

HA-R2R contains 16,844 socially grounded instructions across 90 building scans with 910 annotated human models drawn from HAPS 2.0. Instructions average 112 words, with 20-60% human-related content.

Split / Test SetSizeDetails
Training10,819 instructionsTraining split
Validation: seen778 instructionsSeen validation environments
Validation: unseen1,839 instructionsUnseen validation environments
Held-out test3,408 instructions18 unseen buildings; emphasis on multi-human routes

Evaluation

All submissions are scored server-side on the withheld test split. Test-split ground truth is not released.

Navigation accuracy and social compliance are reported separately. Success requires stopping within 3 m of the goal and completing the episode without collisions. Submissions are ranked by the composite Score defined below.

MetricPreferred DirectionDefinition
Navigation Error (NE)LowerFinal navigation error relative to the goal.
Success Rate (SR)HigherEpisodes that stop within 3 m of the goal and finish without collisions.
Total Collision Rate (TCR)LowerFrequency of collisions in human-occupied zones.
Collision Rate (CR)LowerFraction of human-influenced episodes with at least one collision.

Composite Score

The composite Score uses the full-precision metrics, with higher scores indicating better performance:

Navigation = 0.80×SR + 0.20× 3 3+NE Social = 0.75× (1−CR) + 0.25× 1 1+TCR Score = 100×Navigation× (0.70+0.30×Social)

Higher Score ranks first. Exact ties use higher SR, lower NE, lower CR, lower TCR, then earlier submission time. Calculation and comparison retain full precision.

Reference Baselines

HA-VLN-CMA fuses BERT instruction embeddings with ResNet RGB-D features through cross-modal attention and decodes actions using a GRU policy.

Training uses environmental dropout and DAgger to improve re-planning under partial observability and unpredictable human motion.

Track Organizer

Contact

For challenge questions, contact roboworld2026@gmail.com.