← 피드로
[Submitted on 3 Jun 2024 (v1), last revised 17 Jul 2026 (this version, v5)]
Authors:Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu
Abstract:We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at this https URL.
Submission history
From: Haozheng Luo [view email]
[v1]
Mon, 3 Jun 2024 16:46:18 UTC (2,711 KB)
[v2]
Tue, 4 Jun 2024 03:04:09 UTC (2,711 KB)
[v3]
Thu, 6 Jun 2024 04:25:40 UTC (3,980 KB)
[v4]
Wed, 15 Jul 2026 20:24:21 UTC (4,302 KB)
[v5]
Fri, 17 Jul 2026 07:41:11 UTC (4,301 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2406.01514
답글 남기기