← 피드로
[Submitted on 13 May 2026 (v1), last revised 26 Sep 2026 (this version, v4)]
Abstract:Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervision on hard grounding boundaries, where preferred and rejected responses are semantically close but differ in their support from observable visual evidence. Under a frozen visual encoder and cross-modal alignment pathway, we first use supervised fine-tuning (SFT) to establish broad decoder-side multimodal behavior and then construct hard grounding preference pairs for Direct Preference Optimization (DPO). These pairs target evidence utilization, calibration, and grounding consistency across fine-grained recognition, spatial reasoning, OCR, ambiguous or insufficient evidence, and false-premise queries. The resulting DPO model consistently improves over its SFT initialization on DocVQA, TextVQA, MMBench, and VQAv2. Across approximately 839 additional multimodal examples, it further achieves 6.2–8.2 percentage-point higher pairwise win rates under three independent LLM judges. We additionally release HardVQA-DPO, a curated and growing resource with more than 3K hard-grounding SFT examples and an initial set of DPO preference this http URL://huggingface.co/datasets/vlmgrounding/hardvqa-sft-dpo These results show that decoder-side adaptation alone can improve a meaningful subset of grounding failures under a fixed multimodal representation, while preserving broader multimodal capabilities.
Submission history
From: Qinwu Xu [view email]
[v1]
Wed, 13 May 2026 15:37:51 UTC (3,496 KB)
[v2]
Thu, 6 Aug 2026 13:01:18 UTC (3,452 KB)
[v3]
Tue, 11 Aug 2026 16:06:17 UTC (28,234 KB)
[v4]
Sat, 26 Sep 2026 16:27:15 UTC (35,589 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.16411