Hard Grounding Preference Supervision을 통한 Multimodal Large Language Model의 환각 감소

작성자

카테고리:

← 피드로
arXiv cs.AI · Qinwu Xu · 2026-09-29 AI

[Submitted on 13 May 2026 (v1), last revised 26 Sep 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervision on hard grounding boundaries, where preferred and rejected responses are semantically close but differ in their support from observable visual evidence. Under a frozen visual encoder and cross-modal alignment pathway, we first use supervised fine-tuning (SFT) to establish broad decoder-side multimodal behavior and then construct hard grounding preference pairs for Direct Preference Optimization (DPO). These pairs target evidence utilization, calibration, and grounding consistency across fine-grained recognition, spatial reasoning, OCR, ambiguous or insufficient evidence, and false-premise queries. The resulting DPO model consistently improves over its SFT initialization on DocVQA, TextVQA, MMBench, and VQAv2. Across approximately 839 additional multimodal examples, it further achieves 6.2–8.2 percentage-point higher pairwise win rates under three independent LLM judges. We additionally release HardVQA-DPO, a curated and growing resource with more than 3K hard-grounding SFT examples and an initial set of DPO preference this http URL://huggingface.co/datasets/vlmgrounding/hardvqa-sft-dpo These results show that decoder-side adaptation alone can improve a meaningful subset of grounding failures under a fixed multimodal representation, while preserving broader multimodal capabilities.

Submission history

From: Qinwu Xu [view email]
[v1] Wed, 13 May 2026 15:37:51 UTC (3,496 KB)
[v2] Thu, 6 Aug 2026 13:01:18 UTC (3,452 KB)
[v3] Tue, 11 Aug 2026 16:06:17 UTC (28,234 KB)
[v4] Sat, 26 Sep 2026 16:27:15 UTC (35,589 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.16411