DriveHierarchy: Open-Loop 이해에서 Closed-Loop 실행에 이르기까지 VLM 구동 기능 진단을 위한 벤치마크

작성자

카테고리:

← 피드로
arXiv cs.AI · Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun · 2026-09-29 AI

[Submitted on 25 Sep 2026]

View PDF HTML (experimental)

Abstract:Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on this https URL

Submission history

From: Chengkai Xu [view email]
[v1] Fri, 25 Sep 2026 13:57:44 UTC (31,498 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.31814