← 피드로
[Submitted on 12 Apr 2026 (v1), last revised 7 Sep 2026 (this version, v2)]
Authors:Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao
Abstract:Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is “multi-image reasoning hallucination”, where a large performance drop between forward and reverse temporal queries reveals a dependence on superficial shortcuts instead of state-based understanding. To mitigate this, we first develop a new Chain-of-Thought (CoT) dataset that decomposes intricate reasoning into detailed spatiotemporal steps and definitive judgments. Building on this, we present a progressive training framework: it initiates with supervised pre-training on our CoT dataset to instill logical structures, followed by fine-tuning with scalable weakly-labeled data for broader generalization. Our experiments demonstrate that this approach not only improves backbone accuracy but also reduces the forward-backward performance gap from over 70% to only 6.53%. This shows that the method strengthens dynamic reasoning and reduces the inherent temporal biases of current VLMs.
Submission history
From: Can Wang [view email]
[v1]
Sun, 12 Apr 2026 07:48:44 UTC (2,380 KB)
[v2]
Mon, 7 Sep 2026 12:43:43 UTC (2,129 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.10506