시공간적 환각을 완화하기 위한 구현된 비전 언어 모델을 위한 점진적 훈련 전략

작성자

카테고리:

← 피드로
arXiv cs.AI · Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao · 2026-09-09 AI

[Submitted on 12 Apr 2026 (v1), last revised 7 Sep 2026 (this version, v2)]

Authors:Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao

View PDF HTML (experimental)

Abstract:Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is “multi-image reasoning hallucination”, where a large performance drop between forward and reverse temporal queries reveals a dependence on superficial shortcuts instead of state-based understanding. To mitigate this, we first develop a new Chain-of-Thought (CoT) dataset that decomposes intricate reasoning into detailed spatiotemporal steps and definitive judgments. Building on this, we present a progressive training framework: it initiates with supervised pre-training on our CoT dataset to instill logical structures, followed by fine-tuning with scalable weakly-labeled data for broader generalization. Our experiments demonstrate that this approach not only improves backbone accuracy but also reduces the forward-backward performance gap from over 70% to only 6.53%. This shows that the method strengthens dynamic reasoning and reduces the inherent temporal biases of current VLMs.

Submission history

From: Can Wang [view email]
[v1] Sun, 12 Apr 2026 07:48:44 UTC (2,380 KB)
[v2] Mon, 7 Sep 2026 12:43:43 UTC (2,129 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.10506