← 피드로
[Submitted on 16 Apr 2026 (v1), last revised 31 Aug 2026 (this version, v2)]
Abstract:Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acquired during pre-training. Since these errors arise as a by-product of knowledge degradation, we explore whether established continual learning tools can mitigate them. We propose a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t.~pre-existing knowledge by regularizing output-distribution drift. We also show that when new knowledge acquisition is unnecessary, suppressing factual plasticity by freezing parameter groups preserves task performance while reducing hallucinations. Lastly, we investigate the mechanism, contrasting capacity limitations, behavior cloning, and localized interference. Our experiments show that a main driver is interference among overlapping semantic representations, which self-distillation mitigates and an associative-memory model explains: forgetting grows with the overlap between new and stored facts.
Submission history
From: Guy Kaplan [view email]
[v1]
Thu, 16 Apr 2026 23:08:18 UTC (4,041 KB)
[v2]
Mon, 31 Aug 2026 22:46:31 UTC (4,224 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.15574