← 피드로
[Submitted on 3 Dec 2025 (v1), last revised 14 Jul 2026 (this version, v3)]
Abstract:Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies); these side-effects are commonly referred to as the ripple effect. We introduce RippleBench-Maker, an automatic pipeline that retrieves semantic neighbors of any source concept from a knowledge repository and generates multiple-choice questions at varying semantic distances. We instantiate this framework using WikiRAG, an open-source RAG system over English Wikipedia, to construct RippleBench-WMDP-Bio (584 seed topics, 352,961 questions), and evaluate eight unlearning methods on Llama3-8B-Instruct. All eight exhibit accuracy drops that are largest near the unlearned target and decay with semantic distance, each with a distinct propagation profile. We replicate these findings across Mistral-7B, Zephyr-7B, and Yi-34B; cross-model delta curves are nearly identical, suggesting ripple effects are a property of the unlearning method rather than the base model. We validate all major pipeline stages using a four-experiment Mechanical Turk study (5,200+ responses, 61 workers). We release all code, data, and infrastructure.
Submission history
From: Roy Rinberg [view email]
[v1]
Wed, 3 Dec 2025 18:57:59 UTC (849 KB)
[v2]
Wed, 17 Jun 2026 03:25:12 UTC (1,665 KB)
[v3]
Tue, 14 Jul 2026 05:13:06 UTC (1,664 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2512.04144
답글 남기기