맥락이 진실을 형성하는 방법: LLM에서 문장 수준의 진실 표현의 기하학적 변환

작성자

카테고리:

← 피드로
arXiv cs.AI · Shivam Adarsh, Maria Maistro, Christina Lioma · 2026-06-09 AI

[Submitted on 10 Jan 2026 (v1), last revised 6 Jun 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change when context is introduced remains unexplored. We study this question by measuring (1) the directional change ($theta$) between the truth vectors with and without context and (2) the relative magnitude of the truth vectors upon adding context. Across four LLMs and four datasets, we find that (1) truth vectors are roughly orthogonal in early layers, converge in middle layers, and may stabilize or continue increasing in later layers; (2) adding context generally increases the truth vector magnitude, i.e., the separation between true and false representations in the activation space is amplified; (3) larger models distinguish relevant from irrelevant context mainly through directional change ($theta$), while smaller models show this distinction through magnitude differences. We also find that context conflicting with parametric knowledge produces larger geometric changes than parametrically aligned context. To the best of our knowledge, this is the first work that provides a geometric characterization of how context transforms the truth vector in the activation space of LLMs.

Submission history

From: Shivam Adarsh [view email]
[v1] Sat, 10 Jan 2026 15:43:26 UTC (21,127 KB)
[v2] Sat, 6 Jun 2026 18:23:06 UTC (2,035 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.06599

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다