← 피드로
[Submitted on 22 Sep 2026]
Abstract:Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.
Submission history
From: Florian Lux [view email]
[v1]
Tue, 22 Sep 2026 14:24:10 UTC (9,986 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.26486