← 피드로
[Submitted on 10 Apr 2025 (v1), last revised 9 Jul 2026 (this version, v4)]
Abstract:Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a one-dimensional view of difficulty. In this work, we reframe alignment difficulty as a two-dimensional space spanned by Prompt Complexity (PC) and Pairwise Distinguishability (PD), providing a more principled foundation for alignment. We first demonstrate the efficacy of this space by developing DM-Curri-DPO, a framework of static curricula that already achieves significant gains over baseline methods. Moving beyond these handcrafted paths, we introduce our primary contribution: GSP-Curri-DPO, a novel Group-wise Self-Paced Learning framework. This advanced method empowers the model to navigate the difficulty grid, discovering an optimal learning trajectory based on its own evolving capabilities. Extensive experiments show our self-paced approach not only sets a new state-of-the-art on key benchmarks but, more importantly, demonstrates superior data efficiency and robustness to preference noise. Our work establishes a new paradigm for LLM alignment, offering both a structured difficulty space and an intelligent, model-driven methodology for navigating it.
Submission history
From: Mengyang Li [view email]
[v1]
Thu, 10 Apr 2025 15:32:00 UTC (1,754 KB)
[v2]
Mon, 21 Apr 2025 01:18:57 UTC (1,187 KB)
[v3]
Tue, 29 Jul 2025 09:21:07 UTC (1 KB) (withdrawn)
[v4]
Thu, 9 Jul 2026 13:57:40 UTC (197 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2504.07856
답글 남기기