RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

작성자

카테고리:

← 피드로
arXiv cs.AI · Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, Rohit Gandikota · 2026-07-15 AI

[Submitted on 3 Dec 2025 (v1), last revised 14 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies); these side-effects are commonly referred to as the ripple effect. We introduce RippleBench-Maker, an automatic pipeline that retrieves semantic neighbors of any source concept from a knowledge repository and generates multiple-choice questions at varying semantic distances. We instantiate this framework using WikiRAG, an open-source RAG system over English Wikipedia, to construct RippleBench-WMDP-Bio (584 seed topics, 352,961 questions), and evaluate eight unlearning methods on Llama3-8B-Instruct. All eight exhibit accuracy drops that are largest near the unlearned target and decay with semantic distance, each with a distinct propagation profile. We replicate these findings across Mistral-7B, Zephyr-7B, and Yi-34B; cross-model delta curves are nearly identical, suggesting ripple effects are a property of the unlearning method rather than the base model. We validate all major pipeline stages using a four-experiment Mechanical Turk study (5,200+ responses, 61 workers). We release all code, data, and infrastructure.

Submission history

From: Roy Rinberg [view email]
[v1] Wed, 3 Dec 2025 18:57:59 UTC (849 KB)
[v2] Wed, 17 Jun 2026 03:25:12 UTC (1,665 KB)
[v3] Tue, 14 Jul 2026 05:13:06 UTC (1,664 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2512.04144

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다