← 피드로
[Submitted on 19 Nov 2025 (v1), last revised 17 Jul 2026 (this version, v4)]
Abstract:Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We introduce DiffuMamba, a masked diffusion language model built on a bidirectional Mamba backbone that combines the diffusion objective with linear-time sequence modeling, and DiffuMamba-H, a hybrid variant with interleaved attention. Across scales up to 1.3B parameters, our models match Transformer-based diffusion in downstream performance while achieving up to 8.2x and 4.3x higher inference throughput, respectively, on long sequences. We further present a systematic analysis of inference efficiency across modern DLM variants combining asymptotic complexity with empirical measurements. Notably, cache-efficient block diffusion with Mamba mixers emerges as the only strategy that scales linearly with sequence length and achieves the strongest performance across all baselines, suggesting a promising direction for future diffusion-based generation systems.
Submission history
From: Vaibhav Singh [view email]
[v1]
Wed, 19 Nov 2025 23:23:49 UTC (240 KB)
[v2]
Sun, 23 Nov 2025 05:32:34 UTC (247 KB)
[v3]
Fri, 27 Feb 2026 12:57:37 UTC (213 KB)
[v4]
Fri, 17 Jul 2026 11:40:17 UTC (420 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2511.15927
답글 남기기