스마트 저전력 청각 장치에서 실시간 음성 향상을 위한 트랜스포머 기반 신경 빔포밍

작성자

카테고리:

← 피드로
arXiv cs.AI · Luca Bompani, Marco Fariselli, Giovanni Oltrecolli, Francesco Conti · 2026-09-29 AI

[Submitted on 27 Sep 2026]

View PDF HTML (experimental)

Abstract:Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15~ms per-frame latency while refreshing a complete set of CNN-derived weights every 564~ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65\%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26~dB, and a wideband PESQ of 3.676 at an average power of 45.9~mW. A speech activity detection (SAD) module (98.5\% accuracy, 0.62~mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16~h all-day target on a 100~mAh battery, with an estimated lifetime of up to $\sim$20~h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.

Submission history

From: Luca Bompani [view email]
[v1] Sun, 27 Sep 2026 16:53:28 UTC (1,188 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.33755