HA-VLN 2.0: 동적 다중 인간 상호 작용이 있는 이산적이고 지속적인 환경에서 인간 인식 탐색을 위한 개방형 벤치마크 및 리더보드

작성자

카테고리:

← 피드로
arXiv cs.AI · Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng · 2026-06-09 AI

[Submitted on 18 Mar 2025 (v1), last revised 7 Jun 2026 (this version, v4)]

Authors:Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng

View PDF HTML (experimental)

Abstract:Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments. We present HA-VLN 2.0, a unified benchmark introducing explicit social-awareness constraints. Our contributions are: (i) a standardized task and metrics capturing both goal accuracy and personal-space adherence; (ii) HAPS 2.0 dataset and simulators modeling multi-human interactions, outdoor contexts, and finer language-motion alignment; (iii) benchmarks on 16,844 socially grounded instructions, revealing sharp performance drops of leading agents under human dynamics and partial observability; and (iv) real-world robot experiments validating sim-to-real transfer, with an open leaderboard enabling transparent comparison. Results show that explicit social modeling improves navigation robustness and reduces collisions, underscoring necessity of human-centric approaches. By releasing datasets, simulators, baselines, and protocols, HA-VLN 2.0 provides a strong foundation for safe, human-aware navigation research.

Submission history

From: Yifei Dong [view email]
[v1] Tue, 18 Mar 2025 13:05:55 UTC (34,425 KB)
[v2] Tue, 27 May 2025 16:53:43 UTC (34,430 KB)
[v3] Thu, 9 Oct 2025 18:17:24 UTC (12,288 KB)
[v4] Sun, 7 Jun 2026 15:06:46 UTC (14,092 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2503.14229

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다