주의 헤드에서 폐쇄 검증 회로 발견: 공동 활성화 제안, 제거 폐기

작성자

카테고리:

← 피드로
arXiv cs.AI · Yongzhong Xu · 2026-06-09 AI

[Submitted on 8 Jun 2026]

View PDF HTML (experimental)

Abstract:Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask whether such a cheap signal actually identifies an attention-head circuit. Adapting a sparse-autoencoder clustering recipe to attention heads — but validating by causal ablation rather than reconstruction — we cluster heads and then run a closure test: ablate the discovered community and compare per-example damage to matched-random controls. Across two dense 1B-scale models (Pythia 1B, OLMo 1B) and two input distributions, the communities pass closure. In a Mixture-of-Experts model (OLMoE-1B-7B), route-conditional clustering recovers a statistically real signal that nonetheless does not survive closure — ablation improves loss, the wrong direction. Extending closure across training, attention-target selectivity and participation ratio decouple from function in both directions. We conclude that a cheap signal is a circuit proposal, not a confirmed circuit; closure is what separates them.

Submission history

From: Yongzhong Xu [view email]
[v1] Mon, 8 Jun 2026 15:17:54 UTC (107 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.09607

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다