IEEE RA-L

DA-Nav

Direction-Aware City-Scale Vision-Language Navigation

Ye Yuan1,4,5,*, Kehan Chen2,3,*, Xinqiang Yu2,3,*, Wentao Xu1,4, Heng Wang5, Libo Huang4
Chuanguang Yang4, Yan Huang2, Jiawei He5,†, Zhulin An4,†

* equal contribution    † corresponding authors

1 ShanghaiTech University   2 NLPR, Institute of Automation, CAS   3 University of Chinese Academy of Sciences
4 Institute of Computing Technology, CAS   5 XYZ Embodied AI

DA-Nav overview: ReDA data, direction-aware reasoning, and real-world quadruped and humanoid deployment

DA-Nav turns a coarse commercial-map direction into a local target on the egocentric image, then recovers when the robot drifts. The same policy transfers, without real-world fine-tuning, to a quadruped and a humanoid.

Abstract

City-scale outdoor navigation usually depends on dense maps or expensive, finely annotated trajectories. DA-Nav instead uses the sparse directions already produced by commercial navigation tools, such as “turn left in 50 meters.” The policy treats navigation as discrete spatial grounding on the egocentric image: a vision-language model first checks whether the robot has left the path, then chooses a corrective action, and finally selects a short sequence of grid targets. Trained on ReDA, which pairs direction-aware instructions with recovery trajectories, DA-Nav reaches 59.0% success in CARLA and keeps a 98.15% correction success rate. Without fine-tuning, it runs closed-loop on a Unitree Go2 and a Leju Kuavo humanoid for kilometer-scale routes outdoors.
Keywords: Vision-Language Navigation · Trajectory Recovery · Sim-to-Real

ReDA Dataset

A direction-aware training set: the same sparse instructions a phone map already gives, plus the recovery after the robot leaves the path.

ReDA routes over a city map, with egocentric frames, 286K frames, 2102 trajectories, 128K recovery samples, and 126 scenes

ReDA is collected automatically in CARLA. A controller follows the expert route, then a steering perturbation pushes the robot aside. Once the lateral error passes 0.35 m, the logger records the way back and drops the unstable frames in between. Each kept frame carries a chain-of-thought label: whether the robot has left the path, which action corrects it, and the next six targets on the egocentric grid.

The instructions stay coarse on purpose: forward, turn left, turn right, or stop. That is what commercial navigation already outputs, and it is what the policy is trained to ground.

126scenes
2,102trajectories
286Kframes, 158K expert
128Krecovery frames
Dataset Instruction Action space Supervision
TouchdownFine-grained languageGraph nodesExpert only
CityWalkerPath-level3D waypointsExpert only
NaVid / NaVILAPath-level / mixedDiscrete commandsExpert only
ReDA (ours)Direction-based2D image gridExpert + recovery

Method

DA-Nav is a LoRA-finetuned Qwen2.5-VL-7B. It reads the recent view and a direction, then writes a short chain of thought as cells on the image.

DA-Nav architecture: frozen vision encoder and text embedder, Qwen2.5-VL-7B with LoRA, and a chain-of-thought output of state, action, and grid targets
Two seconds of egocentric frames and a direction token. The vision encoder and text embedder stay frozen; LoRA adapts the backbone. The answer is a state, an action, and six grid cells.

1 Deviation

First the model says whether the robot is still on the reference path: yes or no.

2 Action

Then it picks a command: forward, turn left or right, correct left or right, or stop.

3 Grid targets

Last, it selects six cells for the next three seconds, sampled at 2 Hz on the image plane.

Egocentric image grid with predicted trajectory cells, and the same targets lifted into the robot body frame
The ground in front of the robot is a grid, rows 13–23 and columns 0–28. Red cells are the predicted targets. They are lifted into the body frame and tracked by a furthest-point controller, so the slower model step does not twitch the steering.

Real-World Deployment

Zero-shot closed-loop runs. A phone navigation app supplies the direction; the robot only sees egocentric RGB.

Quadruped, Olympic Forest Park

A Unitree Go2 follows live map directions across a crowded plaza and into the park greenway, including long straight segments and upcoming turns.

Humanoid, Park

The same policy, with no real-world fine-tuning, walks a Leju humanoid through an outdoor park route, from open plazas onto the surrounding streets.

Humanoid, Campus

Closed-loop walking along a campus sidewalk and road, holding the route through building frontage and marked turns.

Highlights

59.0%simulation success
98.15%correction success
46.7%real-world success
1.2 kmzero-shot outdoors

Directions, not dense maps

Commercial navigation tools already plan the global route. DA-Nav only consumes a discrete local instruction: forward, turn left, turn right, or stop.

Ground, then recover

The model reads a short image history, decides whether it has drifted, and picks the next target on an egocentric grid instead of regressing 3D waypoints.

One policy, two robots

ReDA supplies both expert and recovery trajectories in CARLA. The resulting policy transfers zero-shot to a quadruped and a humanoid.

Results

Closed-loop comparison on 239 CARLA evaluation routes. CityWalker, ViNT, NaVid, and NaVILA are fine-tuned on the same ReDA training split. Correction success is the fraction of deviations the policy brings back inside the safety margin.

Method SR (%) ↑ SPL ↑ CSR (%) ↑
CityWalker†41.4241.3336.93
ViNT†52.7252.0044.65
NaVid†17.9917.9919.71
NaVILA†22.1822.1815.28
Zero-shot Qwen2.5-VL11.3011.3042.76
DA-Nav (ours)59.0058.6698.15

† Fine-tuned on ReDA, keeping each baseline’s own architecture and objective. Qwen2.5-VL uses its official weights without that fine-tuning.

Outdoors, over 30 trials in streets and parks, DA-Nav reaches 46.7% success and 76.3% route completion. CityWalker and ViNT reach 23.3% and 16.7% success under the same protocol.

BibTeX

@article{yuan2026danav,
  title   = {DA-Nav: Direction-Aware City-Scale Vision-Language Navigation},
  author  = {Yuan, Ye and Chen, Kehan and Yu, Xinqiang and Xu, Wentao and
             Wang, Heng and Huang, Libo and Yang, Chuanguang and Huang, Yan and
             He, Jiawei and An, Zhulin},
  journal = {IEEE Robotics and Automation Letters},
  year    = {2026},
  note    = {arXiv:2607.11638}
}