Dual-Difficulty Curriculum Learning for Direct Preference Optimization

작성자

카테고리:

← 피드로
arXiv cs.AI · Mengyang Li, Haozhan Geng, Zhong Zhang, Shuang Liu · 2026-07-11 AI

[Submitted on 10 Apr 2025 (v1), last revised 9 Jul 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a one-dimensional view of difficulty. In this work, we reframe alignment difficulty as a two-dimensional space spanned by Prompt Complexity (PC) and Pairwise Distinguishability (PD), providing a more principled foundation for alignment. We first demonstrate the efficacy of this space by developing DM-Curri-DPO, a framework of static curricula that already achieves significant gains over baseline methods. Moving beyond these handcrafted paths, we introduce our primary contribution: GSP-Curri-DPO, a novel Group-wise Self-Paced Learning framework. This advanced method empowers the model to navigate the difficulty grid, discovering an optimal learning trajectory based on its own evolving capabilities. Extensive experiments show our self-paced approach not only sets a new state-of-the-art on key benchmarks but, more importantly, demonstrates superior data efficiency and robustness to preference noise. Our work establishes a new paradigm for LLM alignment, offering both a structured difficulty space and an intelligent, model-driven methodology for navigating it.

Submission history

From: Mengyang Li [view email]
[v1] Thu, 10 Apr 2025 15:32:00 UTC (1,754 KB)
[v2] Mon, 21 Apr 2025 01:18:57 UTC (1,187 KB)
[v3] Tue, 29 Jul 2025 09:21:07 UTC (1 KB) (withdrawn)
[v4] Thu, 9 Jul 2026 13:57:40 UTC (197 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2504.07856

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다