← 피드로
[Submitted on 28 Jul 2026]
Abstract:Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model’s unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.
Submission history
From: Chong Yong Ong [view email]
[v1]
Tue, 28 Jul 2026 08:46:06 UTC (3,870 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.11227
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.