Edge Phoneme Recognition for Children's Speech through Age-Aware Training

작성자

카테고리:

← 피드로
arXiv cs.AI · Matthew Arboleda, Ryan Arboleda, Sophie Haak, Sam Hjelmeset, Andrew Franck, Bingrui Yang, Jose Bustamante Ortiz, Yuanrong Shen, Joel Walsh · 2026-08-12 AI

[Submitted on 10 Aug 2026]

View PDF HTML (experimental)

Abstract:Detecting phonemes from children’s speech has historically been difficult due to the scarcity of training data, and unique characteristics of children’s speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children’s speech, with the privacy and compliance benefits that come with edge processing.

Submission history

From: Joel Walsh [view email]
[v1] Mon, 10 Aug 2026 20:20:56 UTC (344 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.10206

코멘트

답글 남기기