← 피드로
[Submitted on 18 Jun 2026 (v1), last revised 13 Jul 2026 (this version, v3)]
Abstract:Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of diffusion transformer blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a novel instruction-guided audio editing framework based on rectified flow matching (RFM), named RFM-Editing 2, built on a hybrid two-stage diffusion transformer. The proposed model performs joint attention over audio and text tokens to establish coarse semantic alignment at the low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at the high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency.
Submission history
From: Liting Gao [view email]
[v1]
Thu, 18 Jun 2026 11:20:08 UTC (14,069 KB)
[v2]
Wed, 1 Jul 2026 21:49:41 UTC (14,069 KB)
[v3]
Mon, 13 Jul 2026 22:06:08 UTC (17,180 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.20101
답글 남기기