라우팅 그래프로서의 주의: 단일 순방향 패스에서 실시간 회로 추출

작성자

카테고리:

← 피드로
arXiv cs.AI · Ash Manvi, Samreena Tajreen · 2026-09-23 AI

[Submitted on 21 Sep 2026]

View PDF HTML (experimental)

Abstract:Finding circuits in language models usually means running many careful interventions. We try something simpler: treat attention as a routing map from one forward pass, keep a small set of routes that point toward the answer, and ask whether those routes actually matter.
They often do. On induction and IOI (tasks where the “right” circuit is already known), ablating our extracted edges hurts the model much more than ablating a random set of the same size. We evaluate n=100 prompts per cell on GPT-2 Small, GPT-2 Medium, and Pythia-410M, with paired gap tests and bootstrap confidence intervals. The extract step costs one forward; a head-by-head patch sweep costs about two orders of magnitude more.
We are not claiming a complete circuit atlas. We are claiming a cheap sketch that carries real causal signal on known tasks, with clear failure modes when it does not. Code and evaluation artifacts are at this https URL.

Submission history

From: Ash Manvi [view email]
[v1] Mon, 21 Sep 2026 18:27:16 UTC (320 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.25285