Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability

작성자

카테고리:

← 피드로
arXiv cs.AI · Xinyan Jiang, Ninghao Liu, Di Wang, Lijie Hu · 2026-06-16 AI

[Submitted on 11 Mar 2026 (v1), last revised 14 Jun 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality through theoretically grounded geometric kinematics. By decomposing reasoning traces into Progress (displacement) and Stability (curvature), we reveal a distinct topological divergence: correct reasoning manifests as high-progress, stable trajectories, whereas hallucinations are characterized by low-progress, unstable patterns (stalled displacement with high curvature fluctuations). Leveraging these signatures, our probabilistic framework achieves competitive performance and superior robustness across diverse benchmarks. Crucially, TRACED bridges geometry and cognition by mapping high curvature to ”Hesitation Loops” and displacement to ”Certainty Accumulation”, offering a physical lens to decode the internal dynamics of machine thought.

Submission history

From: Xinyan Jiang [view email]
[v1] Wed, 11 Mar 2026 03:58:43 UTC (4,786 KB)
[v2] Sun, 3 May 2026 02:28:16 UTC (4,776 KB)
[v3] Sun, 14 Jun 2026 06:48:37 UTC (4,765 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2603.10384

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다