Cultural Binding Heads in Language Models

작성자

카테고리:

← 피드로
arXiv cs.AI · Avrile Floro, Luca Benedetto · 2026-07-07 AI

[Submitted on 27 May 2026 (v1), last revised 4 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (four architectures, base and instruct). Cultural binding is the process of associating cultural items with the appropriate identity. Knockout of the identity-to-item edges on these heads lowers the binding strength by 9-23%. The identified heads transfer from instruct to base models, suggesting that cultural binding is created at pre-training. An $\alpha$-scaling shows a graded dose-response and moderate amplification steering at generation ($\alpha = 2-3$) increases cultural differentiation accuracy by 1-3 pp while leaving neutral reasoning mostly intact. A knowledge probing task shows that models know 3-5 times more than they act upon it, indicating that the bottleneck lies in routing and not knowledge.

Submission history

From: Avrile Floro [view email]
[v1] Wed, 27 May 2026 14:35:42 UTC (56 KB)
[v2] Sat, 4 Jul 2026 14:51:18 UTC (53 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.28543

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다