Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting

작성자

카테고리:

← 피드로
arXiv cs.AI · Loay Mualem, Llu'is Pastor-P'erez, Vinh Tong, Andrei Manolache, Tanja Bien, Steffen Staab, Mathias Niepert · 2026-10-01 AI

[Submitted on 29 Sep 2026]

View PDF HTML (experimental)

Abstract:Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.

Submission history

From: Loay Mualem [view email]
[v1] Tue, 29 Sep 2026 18:42:43 UTC (224 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.38385