Decoupled Alignment for Robust Plug-and-Play Adaptation

작성자

카테고리:

← 피드로
arXiv cs.AI · Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu · 2026-07-20 AI

[Submitted on 3 Jun 2024 (v1), last revised 17 Jul 2026 (this version, v5)]

Authors:Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu

View PDF HTML (experimental)

Abstract:We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at this https URL.

Submission history

From: Haozheng Luo [view email]
[v1] Mon, 3 Jun 2024 16:46:18 UTC (2,711 KB)
[v2] Tue, 4 Jun 2024 03:04:09 UTC (2,711 KB)
[v3] Thu, 6 Jun 2024 04:25:40 UTC (3,980 KB)
[v4] Wed, 15 Jul 2026 20:24:21 UTC (4,302 KB)
[v5] Fri, 17 Jul 2026 07:41:11 UTC (4,301 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2406.01514

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다