ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

작성자

카테고리:

← 피드로
arXiv cs.AI · Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim · 2026-06-23 AI

[Submitted on 4 Mar 2026 (v1), last revised 19 Jun 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality. Audio samples are available on our web page.

Submission history

From: Youngwon Choi [view email]
[v1] Wed, 4 Mar 2026 16:04:02 UTC (1,454 KB)
[v2] Thu, 18 Jun 2026 11:08:47 UTC (1,454 KB)
[v3] Fri, 19 Jun 2026 02:55:38 UTC (1,454 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2603.04219

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다