๐ Originally published (in Japanese) at forge.workstyle.tech.
I received a report about the avatar for the inquiry desk:
“ใใใใใใใใพใกใใ ใใ” becomes “ใใใใใพใกใใ ใใ”.
Since I had just retrained the voice model multiple times, I first suspected the model. To cut to the chase, the model, parameters, and cache were all fine, but the input text passed to TTS was corrupted.
Debugging from the downstream
I’ll list my suspicions in the order I checked them. This order itself is a lesson learned.
Model: I traced the database to see which model the inquiry site was using. It was correctly using the latest trained model.
Cache: There was a TTS cache table, so I checked if it was returning old audio. The entries were from a different provider and over a month old, so they were unrelated.
Synthesis parameters: The runtime was using style_weight=2.0 and sdp_ratio=0.8. My validation used 1.0 / 0.4, so I thought this might be the cause. I tested various combinations:
style_weight 1.0 / 2.0 / 3.0 / 4.0 โ all ratio 1.00
sdp_ratio 0.2 / 0.4 / 0.6 / 0.8 โ all ratio 1.00
All 12 styles ร weight 2.0 โ all ratio 1.00
Enter fullscreen mode Exit fullscreen mode
Emotional styles: I synthesized “ๅฐใ ใๅพ ใกใใ ใใใ” with all 12 styles, and they were all normal.
Pipeline: The voice pipeline had been moved to a separate service, so I checked if it was synthesizing independently. Synthesis was handled on the backend, and the pipeline was the same.
After several hours, everything checked out.
Printing the preprocessed output once
The only thing left was the input text.
>>> _clean_tts_text('ๅฐใ
ใๅพ
ใกใใ ใใใ')
'ๅฐใๅพ
ใกใใ ใใใ'
Enter fullscreen mode Exit fullscreen mode
The repeating character “ใ
” was missing. And when this corrupted text was synthesized, it sounded like this:
Input 'ๅฐใ
ใๅพ
ใกใใ ใใ' โ Heard 'ๅฐใ
ใๅพ
ใกใใ ใใ' (normal)
Input 'ๅฐใๅพ
ใกใใ ใใ' โ Heard 'ใทใงใผใใๅพ
ใกใใ ใใ' โ this
Enter fullscreen mode Exit fullscreen mode
“ใทใงใผใ” was heard as “ใใใ”. A 5-minute check was done last.
Cause: “ใ ” is not in the Kanji range
The preprocessor had a whitelist to remove emoticons, emojis, and special characters.
# TTS preprocessing: Remove unnecessary symbols and emoticons
_TTS_ALLOWED_RE = re.compile(
r"[^ใ-ใ" # Hiragana
r"ใ -ใฟ" # Katakana
r"ไธ-้ฟฟ" # Kanji
r"๏ฝฆ-๏พ" # Half-width Katakana
r"a-zA-Z๏ฝ-๏ฝ๏ผก-๏ผบ"
r"0-9๏ผ-๏ผ"
r"ใใ๏ผ๏ผ,.ใผ"
r"\s"
r"]"
)
Enter fullscreen mode Exit fullscreen mode
The intention is clear, and the implementation is straightforward. The problem is that ไธ-้ฟฟ (CJK Unified Ideographs) does not include “ใ
”. “ใ
” is ใ
, which is in the CJK Symbols and Punctuation block. It’s classified as a symbol, not a Kanji character.
For Japanese speakers, “ใ ” is considered a Kanji character, but Unicode classifies it differently.
Found 7 missing characters
Since one character was missing, there might be others. I checked all characters used in Japanese that are outside the CJK Unified Ideographs.
Character Example Result Impact ใ ๅฐใ ใปๆฅใ ๅฐใปๆฅ “ใใใ” ใ ใๅ ๅ “ใใ” ใ ใๆใๆฅ ๆๆฅ Dates disappear ๐ ฎท (CJK Extension A) ๐ ฎท้ๅฎถ ้ๅฎถ Proper nouns break ้ซ ๏จ (Compatibility Ideographs) ๏จๅฑฑ ๅฑฑ Names break ใ 10ใ20ๅ 1020ๅ Numbers become something else : 3:30 330 Times become something elseAt the inquiry desk, “๏จๅฑฑใใพ” becoming “ๅฑฑใใพ” is quite bad. “10ใ20ๅ” becoming “ใใใซใใ ใฃใทใ” is similarly problematic.
On the other hand, % & ใใ are also removed, but this is intentional as they’re unnecessary for reading. It’s not about keeping everything.
Keeping symbols didn’t fix it
I straightforwardly added ใ and : to the allowlist. It got worse.
'ๅๅพ330ใซ้ๅงใใพใ' โ Heard 'ๅๅพ330ใซโฆ' (numbers incorrect)
'ๅๅพ3:30ใซ้ๅงใใพใ' โ Heard '5ใ30ใ30ใซ้ๅงใใพใ' โ worse when kept
Enter fullscreen mode Exit fullscreen mode
TTS couldn’t interpret : as a time and produced noise like “ใใ”. ใ was treated as a comma, not “ใใ”.
Removing changes the meaning, keeping makes it unreadable. Neither was correct.
Opening up to Japanese
The correct solution was a third option: convert to Japanese at the preprocessing stage.
_TTS_CLEAN_PATTERNS = [
# โ ๏ธ Symbols with numerical meaning are neither removed nor kept but "converted to Japanese".
# Removing turns "10ใ20ๅ" into "1020ๅ", and keeping makes TTS unreadable,
# producing noise like "3:30" โ "5ใ30ใ30" (verified).
(re.compile(r"(\d)\s*[ใ๏ฝ~]\s*(\d)"), r"\1ใใ\2"), # 10ใ20 โ 10ใใ20
(re.compile(r"(\d{1,2})\s*[:๏ผ]\s*(\d{2})"), r"\1ๆ\2ๅ"), # 3:30 โ 3ๆ30ๅ
...
]
Enter fullscreen mode Exit fullscreen mode
โ ๏ธ Place substitutions before removal. Otherwise, 10ใ20 would first become 1020, and the substitution target would be lost.
And add necessary characters to the allowlist.
r"ไธ-้ฟฟ" # Kanji
# โ ๏ธ Characters necessary for Japanese reading but outside CJK Unified Ideographs.
# In practice, "ๅฐใ
ใๅพ
ใกใใ ใใ" became "ๅฐใๅพ
ใกใใ ใใ",
# and was pronounced as "ใทใงใผใใๅพ
ใกใใ ใใ".
r"ใ
ใใป" # ใ
ใ ใป (repeating characters, abbreviations)
r"ใ" # ใ (Chinese numeral zero)
r"ใ-ไถฟ" # CJK Extension A (variant characters, names)
r"่ฑ-๏ซฟ" # CJK Compatibility Ideographs (้ซ ๏จ, etc., name variants)
Enter fullscreen mode Exit fullscreen mode
Verified with actual audio:
'ๅฐใ
ใๅพ
ใกใใ ใใ' โ 'ๅฐใ
ใๅพ
ใกใใ ใใ' โ
'้ซๆฉใป๏จๅฑฑ' โ '้ซๆฉใป๏จๅฑฑ' โ
'10ใใ20ๅใปใฉใใใใพใ' โ '10ใใ20ๅใปใฉใใใใพใ' โ
'ๅๅพ3ๆ30ๅใซ้ๅงใใพใ' โ 'ๅๅพ3ๆ30ๅใซ้ๅงใใพใ' โ
'ๅไปใฏ9ๆ00ๅใใ17ๆ00ๅใงใ' โ '9ๆ0ๅใใ17ๆ0ๅ' โ
Enter fullscreen mode Exit fullscreen mode
โ ๏ธ There’s another pitfall with wave dashes. ใ (U+301C) and ๏ฝ (U+FF5E) are different characters, and which one is used varies by environment. Including only one would miss the other. Both were added.
Also found: Pronunciation dictionary wasn’t applied
During troubleshooting, I discovered that the pronunciation dictionary wasn’t applied at all to this inquiry site. The dictionary is scoped per project, and all 44 existing entries were tied to different projects.
I tested with the inquiry voice:
Notation Correct Pronunciation Actual Pronunciation ๆถฒๅท ใจใญใฌใค ใงใใ ไธปใช ใชใขใ ใชใผใใผ ๅพ้่ชฒ้ ใธใฅใผใชใงใผใซใญใณ ้้ไพกๅค ่กใฃใฆใใพใ ใชใณใใใใคใใน ่จใฃใฆใใพใ“่กใฃใฆใใพใ” becoming “่จใฃใฆใใพใ” is frequent in customer service phrases and changes the meaning.
Of the 44 entries, excluding 10 for company product names and personal names, 34 were general terms needed across all sites (technical terms and Japanese words with split kun’yomi/on’yomi). I deployed these to 3 other projects.
Lessons learned
Check the input first. I spent hours eliminating the model, parameters, cache, and pipeline, only to finish by printing the preprocessed output once. The order was backward.
A whitelist decides what to remove, not what to allow. If you list what to remove, unexpected characters pass through. Listing what to allow turns oversights into immediate omissions. For languages with many character types like Japanese, whitelisting everything is difficult.
When it seems like a remove-or-keep choice, there’s a third option. Symbols were caught between “removing changes meaning” and “keeping makes unreadable,” but there was the option to convert to Japanese. I was stuck thinking preprocessing was “where unnecessary things are removed,” not “where meaning is preserved while form is changed.”
Check all characters in the same category. When “ใ ” was found missing, I investigated other characters that might be missing for the same reason. Seven were found. If I’d stopped at fixing one, the next report would’ve been about name variants.
Series: Mass-producing practical voices from diffusion TTS
This is a record of designing voices from a single caption line, creating training corpora, and mass-producing role-specific practical voices. This article is Part 3: Quality Gates.
All 18 articles in the series
- The TTS chosen for sound quality was too slow for conversation
- Drawing voices like a gacha
- Having a machine select “narrator-like voices” from 24 candidates
- The stricter the quality gate, the more monotone voices survive
- Speaking speed can’t be changed after training
- TTS that changes “recording location” every time it generates
- One rough clip makes the entire style hoarse
- Where did AI’s habit of elongating “ใใใซใกใใผ” come from? 9. “ๅฐใ ” becoming “ใใใ” โ The allowlist was cutting Japanese characters โ You are here
- The hallucination countermeasure code only worked when there was no hallucination
- The “3 characters” allowed by the quality gate became the model’s catchphrase
- I was discarding candidates over fixable flaws
- There are flaws transcription can’t find
- 70 minutes of training material disappeared in a network blink
- How writing “ja” as “JP” created a jargon model
- 4 registration paths, 0 management screens
- Each deployment was overwriting the other’s work
- Pushing unmeasured things with thresholds always fails
The insights are summarized in the notebook on mass-producing practical voices from diffusion TTS.