← 피드로
[Submitted on 9 Jul 2026 (v1), last revised 7 Aug 2026 (this version, v2)]
Abstract:Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3–20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.
Submission history
From: Hantao Zhang [view email]
[v1]
Thu, 9 Jul 2026 22:22:42 UTC (13,767 KB)
[v2]
Fri, 7 Aug 2026 03:49:20 UTC (13,663 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.08970
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.