ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule

작성자

카테고리:

← 피드로
arXiv cs.AI · Yilie Huang, Wenpin Tang, Xunyu Zhou · 2026-10-01 AI

[Submitted on 26 Jan 2026 (v1), last revised 30 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of time steps. We introduce Adaptive Reparameterized Time (ART), which controls the clock speed of a reparameterized time variable to redistribute computation along the sampling trajectory while preserving the terminal time, with the objective of minimizing the aggregate Euler discretization error. We derive a randomized companion ART-RL that recasts ART as a continuous-time reinforcement learning problem with Gaussian policies, and prove a two-directional bridge between the two: the deterministic ART optimum lifts to an optimal Gaussian policy, and conversely any optimal Gaussian policy must recover the ART control through its mean. This bridge turns continuous-time actor–critic learning into a principled, rather than heuristic, route to the deterministic timestep optimum. Within the official EDM pipeline, ART-RL improves FID on CIFAR–10 across a wide range of budgets; after one-time offline training, the distilled deterministic schedule transfers without retraining to AFHQv2, FFHQ, and ImageNet at no extra inference cost.

Submission history

From: Yilie Huang [view email]
[v1] Mon, 26 Jan 2026 16:56:40 UTC (3,166 KB)
[v2] Fri, 8 May 2026 00:14:56 UTC (5,024 KB)
[v3] Wed, 30 Sep 2026 08:37:13 UTC (1,663 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.18681