Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models

작성자

카테고리:

← 피드로
arXiv cs.AI · Yuxiang Chen, Zuohan Wu, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen · 2026-10-01 AI

[Submitted on 30 Nov 2025 (v1), last revised 30 Sep 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise atomic reasoning steps and analyse the reasoning behaviours of LRMs. Grounded in human cognitive processes, we propose a taxonomy comprising five groups and seventeen categories. Through this taxonomy, we conduct an in-depth analysis of contemporary LRMs and distil four actionable takeaways for model optimisation. Most notably, we reveal that prevailing post-answer “doublechecks” are largely superficial and rarely yield substantive revisions. A targeted intervention further shows that explicitly eliciting richer reflection processes can substantially improve failed self-correction. To support this largescale study, we propose CAPO, an automated annotation method used to construct a dataset of 277,534 reasoning steps with strong agreement with human expert annotations. We further validate the main behavioural patterns on a newer reasoning model and a coding domain, demonstrating the broader applicability of the proposed taxonomy. All source code and data are available at this https URL.

Submission history

From: Zuohan Wu [view email]
[v1] Sun, 30 Nov 2025 04:49:44 UTC (1,235 KB)
[v2] Wed, 30 Sep 2026 03:46:25 UTC (1,147 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2512.00729