Forecasting Side Effects of Activation Steering

작성자

카테고리:

← 피드로
arXiv cs.AI · Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun · 2026-08-13 AI

[Submitted on 28 Jul 2026]

View PDF HTML (experimental)

Abstract:Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model’s unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.

Submission history

From: Chong Yong Ong [view email]
[v1] Tue, 28 Jul 2026 08:46:06 UTC (3,870 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.11227

코멘트

답글 남기기