DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

작성자

카테고리:

← 피드로
arXiv cs.AI · Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr'e No"el, Eugene Belilovsky, Torsten Scholak · 2026-07-20 AI

[Submitted on 19 Nov 2025 (v1), last revised 17 Jul 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We introduce DiffuMamba, a masked diffusion language model built on a bidirectional Mamba backbone that combines the diffusion objective with linear-time sequence modeling, and DiffuMamba-H, a hybrid variant with interleaved attention. Across scales up to 1.3B parameters, our models match Transformer-based diffusion in downstream performance while achieving up to 8.2x and 4.3x higher inference throughput, respectively, on long sequences. We further present a systematic analysis of inference efficiency across modern DLM variants combining asymptotic complexity with empirical measurements. Notably, cache-efficient block diffusion with Mamba mixers emerges as the only strategy that scales linearly with sequence length and achieves the strongest performance across all baselines, suggesting a promising direction for future diffusion-based generation systems.

Submission history

From: Vaibhav Singh [view email]
[v1] Wed, 19 Nov 2025 23:23:49 UTC (240 KB)
[v2] Sun, 23 Nov 2025 05:32:34 UTC (247 KB)
[v3] Fri, 27 Feb 2026 12:57:37 UTC (213 KB)
[v4] Fri, 17 Jul 2026 11:40:17 UTC (420 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2511.15927

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다