The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models

작성자

카테고리:

← 피드로
arXiv cs.AI · Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin, Bao Li, Jung-Hsien Chiang · 2026-08-03 AI

[Submitted on 4 Jun 2026 (v1), last revised 31 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources. We ask whether this reflects a capability deficit or an artifact of the role labeling. To test this, we design a training-free intervention, source-conditioned role relabeling, that keeps the erroneous claim byte-identical and varies only its message role. The claim is presented inside the agent’s “<thought>”, a user message, a tool response, or a system “<memory>” block. We test 12 model-domain combinations spanning closed-weight APIs and open-weight models from 70B-class down to smaller families. Relabeling “<thought>” to an external role increases the explicit-correction rate by 23 to 93 percentage points, significant in 10 of 12 experimental settings. This suggests that these models’ failure to detect a self-generated error is largely an artifact of how the claim is role-labeled in the chat template, rather than a pure cognitive deficit. The most effective role label is domain-dependent: “<memory>” dominates in most math experiments, while a user message dominates in logical deduction. Recognizing role-label handling as a key experimental variable in instruction tuning presents a more direct path to closing the self-correction gap.d

Submission history

From: Fang-Yi Su [view email]
[v1] Thu, 4 Jun 2026 10:17:00 UTC (393 KB)
[v2] Fri, 31 Jul 2026 08:26:14 UTC (479 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.05976

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다