ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

작성자

카테고리:

← 피드로
arXiv cs.AI · Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu · 2026-08-03 AI

[Submitted on 23 Jul 2026 (v1), last revised 31 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

Submission history

From: Chi Zhang [view email]
[v1] Thu, 23 Jul 2026 17:15:30 UTC (13,459 KB)
[v2] Fri, 31 Jul 2026 09:23:37 UTC (26,960 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.21529

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다