중단 단계: 재귀 훈련은 재생된 샘플링 노이즈와 공명합니다

작성자

카테고리:

← 피드로
arXiv cs.AI · Yangze Liu, Zhongyi Han · 2026-09-29 AI

[Submitted on 10 Sep 2026 (v1), last revised 29 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments report repeated phrases within ten generations. Under a fixed sampling seed in vLLM, the fast loss of lexical diversity comes from the sampler. When vLLM serves a batch from one seeded sampling configuration, every request receives the same random draws, and a fixed seed replays them every generation. Fine-tuning raises the tokens that won, and the replayed draws let them win by more. Sharing across requests and replay across generations matter only together. Remove either one, by changing the shared seed every generation or by giving each request its own seed that repeats every generation, and the unique-4-gram fraction of two StableLM checkpoints stays near its starting value of about 0.98 through generation 3. Keep both, and the replayed shared seed takes seven checkpoints from five families to between 0.045 and 0.38 by then. Three generations of replay write the favoured phrases into the weights: decoded with one seed per request, the generation-3 weights of the replayed StableLM-2-1.6B chain recover most of their diversity, yet the phrase that filled every sample under the shared seed still opens 46% of them. Without replay, five checkpoints drift slowly, consistent with the gradual accumulation that theory describes, and three turn incoherent though their diversity scores stay high. One peer-reviewed model-collapse pipeline that fine-tunes Gemma-2-27B samples identical prompts under one seeded configuration, and three quarters of the rows it released for one iteration repeat nearly as often as one such batch copies them. A seed per request restores the fresh sample that stability analyses assume.

Submission history

From: Yangze Liu [view email]
[v1] Thu, 10 Sep 2026 06:56:22 UTC (162 KB)
[v2] Sat, 26 Sep 2026 12:15:10 UTC (1,444 KB)
[v3] Tue, 29 Sep 2026 05:48:04 UTC (1,442 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.11149