← 피드로
[Submitted on 8 Jan 2026 (v1), last revised 31 Aug 2026 (this version, v5)]
Abstract:Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rate acoustic features directly into the LLM input space, making long-context processing computationally prohibitive. Unlike images or videos, speech lacks spatial redundancy, making extreme token compression particularly challenging. To address this limitation, we propose FastSLM, a token-efficient architecture featuring the Hierarchical Temporal Abstractor (HTA), which progressively distills acoustic features across multiple temporal scales. HTA achieves an extreme compression rate of 1.67 tokens per second (97% reduction) while preserving essential linguistic information for downstream speech-language understanding. Experimental results demonstrate that FastSLM achieves competitive performance across diverse speech-language tasks while requiring substantially fewer speech tokens and FLOPs than existing speech-language models. The source code and model checkpoints are available at this https URL.
Submission history
From: Junseok Lee [view email]
[v1]
Thu, 8 Jan 2026 07:46:03 UTC (1,898 KB)
[v2]
Mon, 2 Feb 2026 06:22:57 UTC (1,910 KB)
[v3]
Mon, 1 Jun 2026 03:39:22 UTC (1,861 KB)
[v4]
Fri, 28 Aug 2026 01:40:26 UTC (1,861 KB)
[v5]
Mon, 31 Aug 2026 04:21:44 UTC (1,861 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.06199