← 피드로
[Submitted on 1 Feb 2026 (v1), last revised 24 Sep 2026 (this version, v3)]
Abstract:End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments involving heterogeneous road users and rare safety-critical interactions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning. HERMES employs a foundation-model-assisted annotation pipeline to construct structured Long-Tail Scene Context and Long-Tail Planning Context, capturing hazard-centric scene information, maneuver intent, and risk-aware planning guidance. A Tri-Modal Driving Module then integrates multi-view visual observations, historical ego-motion, and long-tail semantic instructions through intent- and risk-aware conditioning for trajectory generation. Extensive experiments on a large-scale real-world long-tail driving benchmark demonstrate consistent improvements over representative recent baselines in overall planning performance and across diverse safety-critical scenarios. Ablation studies further validate the effectiveness and complementary roles of the major components within HERMES.
Submission history
From: Weizhe Tang [view email]
[v1]
Sun, 1 Feb 2026 03:15:08 UTC (7,888 KB)
[v2]
Fri, 18 Sep 2026 03:00:07 UTC (3,925 KB)
[v3]
Thu, 24 Sep 2026 17:37:00 UTC (3,925 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2602.00993