← 피드로
[Submitted on 22 Apr 2026 (v1), last revised 23 Aug 2026 (this version, v3)]
Abstract:We introduce DialToM, an annotated Theory of Mind (ToM) benchmark built from naturalistic human-human dialogues using a multiple-choice evaluation framework. Concurrent with recent work showing a gap between explicit mental-state inference and applied ToM in synthetic settings~\cite{gu2024simpletom}, we establish a stricter \emph{State-Driven Diagnostic Probe} in which models must forecast state-consistent dialogue trajectories solely from isolated mental-state profiles without dialogue context. Our evaluation reveals a systematic reasoning asymmetry — LLMs excel at inferring mental states (Literal ToM) but struggle to leverage them for social forecasting (Functional ToM). Crucially, a domain expert achieves 100\% accuracy on this task, proving its validity and establishing a stark human-AI capability gap. Further, a teacher-student reasoning injection probe shows that Gemini 3 Pro — which establishes the leading baseline — possesses robust Functional ToM capabilities for context-free forecasting that are transferable to weaker models. DialToM, its evaluation code, and dataset are publicly available at this https URL.
Submission history
From: Palakorn Achananuparp [view email]
[v1]
Wed, 22 Apr 2026 11:07:46 UTC (138 KB)
[v2]
Thu, 28 May 2026 03:44:08 UTC (200 KB)
[v3]
Sun, 23 Aug 2026 08:28:31 UTC (208 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.20443