FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

작성자

카테고리:

← 피드로
arXiv cs.AI · Junseok Lee, Chang-Jae Chun · 2026-08-31 AI

[Submitted on 8 Jan 2026 (v1), last revised 31 Aug 2026 (this version, v5)]

View PDF HTML (experimental)

Abstract:Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rate acoustic features directly into the LLM input space, making long-context processing computationally prohibitive. Unlike images or videos, speech lacks spatial redundancy, making extreme token compression particularly challenging. To address this limitation, we propose FastSLM, a token-efficient architecture featuring the Hierarchical Temporal Abstractor (HTA), which progressively distills acoustic features across multiple temporal scales. HTA achieves an extreme compression rate of 1.67 tokens per second (97% reduction) while preserving essential linguistic information for downstream speech-language understanding. Experimental results demonstrate that FastSLM achieves competitive performance across diverse speech-language tasks while requiring substantially fewer speech tokens and FLOPs than existing speech-language models. The source code and model checkpoints are available at this https URL.

Submission history

From: Junseok Lee [view email]
[v1] Thu, 8 Jan 2026 07:46:03 UTC (1,898 KB)
[v2] Mon, 2 Feb 2026 06:22:57 UTC (1,910 KB)
[v3] Mon, 1 Jun 2026 03:39:22 UTC (1,861 KB)
[v4] Fri, 28 Aug 2026 01:40:26 UTC (1,861 KB)
[v5] Mon, 31 Aug 2026 04:21:44 UTC (1,861 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.06199