Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

작성자

카테고리:

← 피드로
arXiv cs.AI · Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee · 2026-08-12 AI

[Submitted on 10 Aug 2026]

View PDF HTML (experimental)

Abstract:Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology across two domain granularities, three model families (1.5B to 7B parameters), and eight domains. At the academic subject level, zero neurons exceed 60\% domain selectivity across 939,008 combined FFN neurons and causal damage matrices are flat, despite domain identity being linearly decodable above 85\% accuracy. At the language and modality level, 0.65–1.14\% of neurons exceed 60\% selectivity, damage matrices are near-perfectly diagonal (ratios up to 595:1), and shell neuron sets are essentially disjoint (IoU $< 0.003$). Masking code-selective neurons reduces mathematical reasoning accuracy by 16–24 percentage points across all models; masking Spanish or Chinese neurons leaves it at or below random. Shell strength increases monotonically with scale and shells are spatially interleaved in a pattern that precludes group-level selective quantization. Parametric shells form where and only where training data was modular at the token level.

Submission history

From: Marcus Armstrong [view email]
[v1] Mon, 10 Aug 2026 20:35:43 UTC (142 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.10214

코멘트

답글 남기기