Thinking Ahead: Foresight Intelligence in MLLMs and World Model

작성자

카테고리:

← 피드로
arXiv cs.AI · Zhantao Gong, Liaoyuan Fan, Qing Guo, Xun Xu, Xulei Yang, Shijie Li · 2026-07-09 AI

[Submitted on 24 Nov 2025 (v1), last revised 8 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet largely overlooked by existing research. To bridge this gap, we introduce FSU-QA, a new Visual Question-Answering (VQA) dataset specifically designed to elicit and evaluate Foresight Intelligence. Using FSU-QA, we conduct the first comprehensive study of state-of-the-art Vision-Language Models (VLMs) under foresight-oriented tasks, revealing that current models still struggle to reason about future situations. Beyond serving as a benchmark, FSU-QA also enables the assessment of world models by measuring the semantic coherence of their generated predictions, quantified through performance gains when VLMs are augmented with such outputs. Our experiments further demonstrate that FSU-QA can effectively enhance foresight reasoning: even small VLMs fine-tuned on FSU-QA surpass much larger, advanced models by a substantial margin. Together, these findings position FSU-QA as a principled foundation for developing next-generation models capable of truly anticipating and understanding future events. Furthermore, beyond model performance, we examine whether WM-generated predictions remain semantically consistent by using VLM-based proxy judges, and validate this evaluation protocol through shuffled control experiments. Fine-tuning models on FSU-QA leads to substantial improvements in foresight understanding, demonstrating the dataset’s effectiveness and offering a principled foundation for future research.

Submission history

From: Zhantao Gong [view email]
[v1] Mon, 24 Nov 2025 04:04:59 UTC (17,223 KB)
[v2] Thu, 11 Dec 2025 08:02:01 UTC (16,080 KB)
[v3] Wed, 8 Jul 2026 02:31:12 UTC (16,053 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2511.18735

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다