The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

작성자

카테고리:

← 피드로
arXiv cs.AI · Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk · 2026-08-11 AI

[Submitted on 21 Jul 2026]

View PDF HTML (experimental)

Abstract:Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered “persistence beats peak” hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.

Submission history

From: Justin Shenk [view email]
[v1] Tue, 21 Jul 2026 12:13:10 UTC (135 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.07528

코멘트

답글 남기기