Jacobian Scopes: token-level causal attributions in LLMs

작성자

카테고리:

← 피드로
arXiv cs.AI · Toni J. B. Liu, Baran Zadeou{g}lu, Nicolas Boull'e, Rapha"el Sarfati, Gurbir Arora, Christopher J. Earls · 2026-06-15 AI

[Submitted on 23 Jan 2026 (v1), last revised 15 Jun 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Large language models (LLMs) make next-token predictions based on clues present in their context, such as semantic descriptions and in-context examples. Yet, elucidating which prior tokens most strongly influence a given prediction remains challenging due to the proliferation of layers and attention heads in modern architectures. We propose Jacobian Scopes, a suite of gradient-based, token-level causal attribution methods for interpreting LLM predictions. Grounded in perturbation theory and information geometry, Jacobian Scopes quantify how input tokens influence various aspects of a model’s prediction, such as specific logits, the full predictive distribution, and model uncertainty (effective temperature). Through case studies spanning instruction understanding, translation, and in-context learning (ICL), we demonstrate how Jacobian Scopes reveal implicit political biases, uncover word- and phrase-level translation strategies, and shed light on recently debated mechanisms underlying in-context time-series forecasting. To facilitate exploration of Jacobian Scopes on custom text, we open-source our implementations and provide a cloud-hosted interactive demo at this https URL.

Submission history

From: Jianbang Liu [view email]
[v1] Fri, 23 Jan 2026 02:36:38 UTC (7,914 KB)
[v2] Sun, 15 Mar 2026 23:40:50 UTC (9,356 KB)
[v3] Thu, 11 Jun 2026 18:04:05 UTC (9,440 KB)
[v4] Mon, 15 Jun 2026 18:41:08 UTC (9,441 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.16407

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다