SUAN: 대규모 언어 모델에서 직접 기본 설정 안전 정렬 수정

작성자

카테고리:

← 피드로
arXiv cs.AI · Oleksandr Cherednichenko, Roman Klypa · 2026-09-09 AI

[Submitted on 8 Sep 2026 (v1), last revised 2 Oct 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.

Submission history

From: Oleksandr Cherednichenko [view email]
[v1] Tue, 8 Sep 2026 12:04:38 UTC (426 KB)
[v2] Fri, 2 Oct 2026 10:27:05 UTC (426 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.08634