MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

작성자

카테고리:

← 피드로
arXiv cs.AI · Hantao Zhang, Jinru Sui, Ed Li, Dirk Bergemann, Zhuoran Yang · 2026-08-10 AI

[Submitted on 9 Jul 2026 (v1), last revised 7 Aug 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3–20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.

Submission history

From: Hantao Zhang [view email]
[v1] Thu, 9 Jul 2026 22:22:42 UTC (13,767 KB)
[v2] Fri, 7 Aug 2026 03:49:20 UTC (13,663 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.08970

코멘트

답글 남기기