← 피드로
[Submitted on 24 Jul 2025 (v1), last revised 24 Aug 2026 (this version, v4)]
Authors:Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li
Abstract:Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interactional behavior grounded in acoustic context. To address this gap, we introduce TELEVAL, a large-scale SLM benchmark for Chinese spoken interaction in instruction-free, audio-conditioned settings. TELEVAL evaluates two complementary aspects: (1) Reliable Content Fulfillment, which measures semantic accuracy of SLMs under diverse acoustic and linguistic conditions, and (2) Interactional Appropriateness, which assesses whether models produce natural and appropriate responses by implicitly grounding behavior in auditory cues. Experiments show that while models perform competitively on semantic tasks, their performance degrades under acoustic variability and in interactional settings. We observe consistent degradation from perceptual instability to interactional errors, and further identify a recurring failure pattern, termed the “Caption Trap”, where models tend to describe perceived audio signals rather than produce appropriate interactive responses. These results indicate that current SLMs remain insufficiently aligned with the requirements of natural spoken interaction. TELEVAL provides a targeted framework for evaluating and analyzing interactional behavior in SLMs.
Submission history
From: Zehan Li [view email]
[v1]
Thu, 24 Jul 2025 03:23:55 UTC (356 KB)
[v2]
Tue, 6 Jan 2026 08:48:39 UTC (533 KB)
[v3]
Mon, 12 Jan 2026 03:01:57 UTC (533 KB)
[v4]
Mon, 24 Aug 2026 13:42:31 UTC (542 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2507.18061