How 'Shลshล' Became 'Shomo' โ€” Permission Character List Was Trimming Japanese

์ž‘์„ฑ์ž

์นดํ…Œ๊ณ ๋ฆฌ:

โ† ํ”ผ๋“œ๋กœ
DEV Community ยท orca_forge ยท 2026-09-02 ๊ฐœ๋ฐœ(SW)

orca_forge

๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech.

I received a report about the avatar for the inquiry desk:

“ใ—ใ‚‡ใ†ใ—ใ‚‡ใ†ใŠใพใกใใ ใ•ใ„” becomes “ใ—ใ‚‡ใ‚‚ใŠใพใกใใ ใ•ใ„”.

Since I had just retrained the voice model multiple times, I first suspected the model. To cut to the chase, the model, parameters, and cache were all fine, but the input text passed to TTS was corrupted.

Debugging from the downstream

I’ll list my suspicions in the order I checked them. This order itself is a lesson learned.

Model: I traced the database to see which model the inquiry site was using. It was correctly using the latest trained model.

Cache: There was a TTS cache table, so I checked if it was returning old audio. The entries were from a different provider and over a month old, so they were unrelated.

Synthesis parameters: The runtime was using style_weight=2.0 and sdp_ratio=0.8. My validation used 1.0 / 0.4, so I thought this might be the cause. I tested various combinations:

style_weight 1.0 / 2.0 / 3.0 / 4.0   โ†’ all ratio 1.00
sdp_ratio    0.2 / 0.4 / 0.6 / 0.8   โ†’ all ratio 1.00
All 12 styles ร— weight 2.0             โ†’ all ratio 1.00

Enter fullscreen mode Exit fullscreen mode

Emotional styles: I synthesized “ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„ใ€‚” with all 12 styles, and they were all normal.

Pipeline: The voice pipeline had been moved to a separate service, so I checked if it was synthesizing independently. Synthesis was handled on the backend, and the pipeline was the same.

After several hours, everything checked out.

Printing the preprocessed output once

The only thing left was the input text.

>>> _clean_tts_text('ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„ใ€‚')
'ๅฐ‘ใŠๅพ…ใกใใ ใ•ใ„ใ€‚'

Enter fullscreen mode Exit fullscreen mode

The repeating character “ใ€…” was missing. And when this corrupted text was synthesized, it sounded like this:

Input 'ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„' โ†’ Heard 'ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„' (normal)
Input 'ๅฐ‘ใŠๅพ…ใกใใ ใ•ใ„'    โ†’ Heard 'ใ‚ทใƒงใƒผใ‚’ใŠๅพ…ใกใใ ใ•ใ„'   โ† this

Enter fullscreen mode Exit fullscreen mode

“ใ‚ทใƒงใƒผใ‚’” was heard as “ใ—ใ‚‡ใ‚‚”. A 5-minute check was done last.

Cause: “ใ€…” is not in the Kanji range

The preprocessor had a whitelist to remove emoticons, emojis, and special characters.

# TTS preprocessing: Remove unnecessary symbols and emoticons
_TTS_ALLOWED_RE = re.compile(
    r"[^ใ€-ใ‚Ÿ"   # Hiragana
    r"ใ‚ -ใƒฟ"     # Katakana
    r"ไธ€-้ฟฟ"     # Kanji
    r"๏ฝฆ-๏พŸ"     # Half-width Katakana
    r"a-zA-Z๏ฝ-๏ฝš๏ผก-๏ผบ"
    r"0-9๏ผ-๏ผ™"
    r"ใ€ใ€‚๏ผ๏ผŸ,.ใƒผ"
    r"\s"
    r"]"
)

Enter fullscreen mode Exit fullscreen mode

The intention is clear, and the implementation is straightforward. The problem is that ไธ€-้ฟฟ (CJK Unified Ideographs) does not include “ใ€…”. “ใ€…” is ใ€…, which is in the CJK Symbols and Punctuation block. It’s classified as a symbol, not a Kanji character.

For Japanese speakers, “ใ€…” is considered a Kanji character, but Unicode classifies it differently.

Found 7 missing characters

Since one character was missing, there might be others. I checked all characters used in Japanese that are outside the CJK Unified Ideographs.

Character Example Result Impact ใ€… ๅฐ‘ใ€…ใƒปๆ—ฅใ€… ๅฐ‘ใƒปๆ—ฅ “ใ—ใ‚‡ใ‚‚” ใ€† ใ€†ๅˆ‡ ๅˆ‡ “ใใ‚Š” ใ€‡ ใ€‡ๆœˆใ€‡ๆ—ฅ ๆœˆๆ—ฅ Dates disappear ๐ ฎท (CJK Extension A) ๐ ฎท้‡Žๅฎถ ้‡Žๅฎถ Proper nouns break ้ซ™ ๏จ‘ (Compatibility Ideographs) ๏จ‘ๅฑฑ ๅฑฑ Names break ใ€œ 10ใ€œ20ๅˆ† 1020ๅˆ† Numbers become something else : 3:30 330 Times become something else

At the inquiry desk, “๏จ‘ๅฑฑใ•ใพ” becoming “ๅฑฑใ•ใพ” is quite bad. “10ใ€œ20ๅˆ†” becoming “ใ›ใ‚“ใซใ˜ใ‚…ใฃใทใ‚“” is similarly problematic.

On the other hand, % & ใ€Œใ€ are also removed, but this is intentional as they’re unnecessary for reading. It’s not about keeping everything.

Keeping symbols didn’t fix it

I straightforwardly added ใ€œ and : to the allowlist. It got worse.

'ๅˆๅพŒ330ใซ้–‹ๅง‹ใ—ใพใ™'   โ†’ Heard 'ๅˆๅพŒ330ใซโ€ฆ'          (numbers incorrect)
'ๅˆๅพŒ3:30ใซ้–‹ๅง‹ใ—ใพใ™'  โ†’ Heard '5ใ‚‚30ใ€30ใซ้–‹ๅง‹ใ—ใพใ™'  โ† worse when kept

Enter fullscreen mode Exit fullscreen mode

TTS couldn’t interpret : as a time and produced noise like “ใ”ใ‚‚”. ใ€œ was treated as a comma, not “ใ‹ใ‚‰”.

Removing changes the meaning, keeping makes it unreadable. Neither was correct.

Opening up to Japanese

The correct solution was a third option: convert to Japanese at the preprocessing stage.

_TTS_CLEAN_PATTERNS = [
    # โš ๏ธ Symbols with numerical meaning are neither removed nor kept but "converted to Japanese".
    # Removing turns "10ใ€œ20ๅˆ†" into "1020ๅˆ†", and keeping makes TTS unreadable,
    # producing noise like "3:30" โ†’ "5ใ‚‚30ใ€30" (verified).
    (re.compile(r"(\d)\s*[ใ€œ๏ฝž~]\s*(\d)"), r"\1ใ‹ใ‚‰\2"),   # 10ใ€œ20 โ†’ 10ใ‹ใ‚‰20
    (re.compile(r"(\d{1,2})\s*[:๏ผš]\s*(\d{2})"), r"\1ๆ™‚\2ๅˆ†"),      # 3:30 โ†’ 3ๆ™‚30ๅˆ†
    ...
]

Enter fullscreen mode Exit fullscreen mode

โš ๏ธ Place substitutions before removal. Otherwise, 10ใ€œ20 would first become 1020, and the substitution target would be lost.

And add necessary characters to the allowlist.

r"ไธ€-้ฟฟ"     # Kanji
# โš ๏ธ Characters necessary for Japanese reading but outside CJK Unified Ideographs.
# In practice, "ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„" became "ๅฐ‘ใŠๅพ…ใกใใ ใ•ใ„",
# and was pronounced as "ใ‚ทใƒงใƒผใ‚’ใŠๅพ…ใกใใ ใ•ใ„".
r"ใ€…ใ€†ใ€ป"  # ใ€… ใ€† ใ€ป (repeating characters, abbreviations)
r"ใ€‡"           # ใ€‡ (Chinese numeral zero)
r"ใ€-ไถฟ"    # CJK Extension A (variant characters, names)
r"่ฑˆ-๏ซฟ"    # CJK Compatibility Ideographs (้ซ™ ๏จ‘, etc., name variants)

Enter fullscreen mode Exit fullscreen mode

Verified with actual audio:

'ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„'          โ†’ 'ๅฐ‘ใ€…ใŠๅพ…ใกใใ ใ•ใ„'       โœ…
'้ซ™ๆฉ‹ใƒป๏จ‘ๅฑฑ'                 โ†’ '้ซ™ๆฉ‹ใƒป๏จ‘ๅฑฑ'              โœ…
'10ใ‹ใ‚‰20ๅˆ†ใปใฉใ‹ใ‹ใ‚Šใพใ™'     โ†’ '10ใ‹ใ‚‰20ๅˆ†ใปใฉใ‹ใ‹ใ‚Šใพใ™'  โœ…
'ๅˆๅพŒ3ๆ™‚30ๅˆ†ใซ้–‹ๅง‹ใ—ใพใ™'      โ†’ 'ๅˆๅพŒ3ๆ™‚30ๅˆ†ใซ้–‹ๅง‹ใ—ใพใ™'   โœ…
'ๅ—ไป˜ใฏ9ๆ™‚00ๅˆ†ใ‹ใ‚‰17ๆ™‚00ๅˆ†ใงใ™' โ†’ '9ๆ™‚0ๅˆ†ใ‹ใ‚‰17ๆ™‚0ๅˆ†'        โœ…

Enter fullscreen mode Exit fullscreen mode

โš ๏ธ There’s another pitfall with wave dashes. ใ€œ (U+301C) and ๏ฝž (U+FF5E) are different characters, and which one is used varies by environment. Including only one would miss the other. Both were added.

Also found: Pronunciation dictionary wasn’t applied

During troubleshooting, I discovered that the pronunciation dictionary wasn’t applied at all to this inquiry site. The dictionary is scoped per project, and all 44 existing entries were tied to different projects.

I tested with the inquiry voice:

Notation Correct Pronunciation Actual Pronunciation ๆถฒๅ†ท ใ‚จใ‚ญใƒฌใ‚ค ใงใใ’ ไธปใช ใ‚ชใƒขใƒŠ ใ‚ชใƒผใƒŠใƒผ ๅพ“้‡่ชฒ้‡‘ ใ‚ธใƒฅใƒผใƒชใƒงใƒผใ‚ซใ‚ญใƒณ ้‡้‡ไพกๅ€ค ่กŒใฃใฆใ„ใพใ™ ใ‚ชใ‚ณใƒŠใƒƒใƒ†ใ‚คใƒžใ‚น ่จ€ใฃใฆใ„ใพใ™

“่กŒใฃใฆใ„ใพใ™” becoming “่จ€ใฃใฆใ„ใพใ™” is frequent in customer service phrases and changes the meaning.

Of the 44 entries, excluding 10 for company product names and personal names, 34 were general terms needed across all sites (technical terms and Japanese words with split kun’yomi/on’yomi). I deployed these to 3 other projects.

Lessons learned

Check the input first. I spent hours eliminating the model, parameters, cache, and pipeline, only to finish by printing the preprocessed output once. The order was backward.

A whitelist decides what to remove, not what to allow. If you list what to remove, unexpected characters pass through. Listing what to allow turns oversights into immediate omissions. For languages with many character types like Japanese, whitelisting everything is difficult.

When it seems like a remove-or-keep choice, there’s a third option. Symbols were caught between “removing changes meaning” and “keeping makes unreadable,” but there was the option to convert to Japanese. I was stuck thinking preprocessing was “where unnecessary things are removed,” not “where meaning is preserved while form is changed.”

Check all characters in the same category. When “ใ€…” was found missing, I investigated other characters that might be missing for the same reason. Seven were found. If I’d stopped at fixing one, the next report would’ve been about name variants.

Series: Mass-producing practical voices from diffusion TTS

This is a record of designing voices from a single caption line, creating training corpora, and mass-producing role-specific practical voices. This article is Part 3: Quality Gates.

All 18 articles in the series

  1. The TTS chosen for sound quality was too slow for conversation
  2. Drawing voices like a gacha
  3. Having a machine select “narrator-like voices” from 24 candidates
  4. The stricter the quality gate, the more monotone voices survive
  5. Speaking speed can’t be changed after training
  6. TTS that changes “recording location” every time it generates
  7. One rough clip makes the entire style hoarse
  8. Where did AI’s habit of elongating “ใ“ใ‚“ใซใกใ‚ใƒผ” come from? 9. “ๅฐ‘ใ€…” becoming “ใ—ใ‚‡ใ‚‚” โ€” The allowlist was cutting Japanese characters โ† You are here
  9. The hallucination countermeasure code only worked when there was no hallucination
  10. The “3 characters” allowed by the quality gate became the model’s catchphrase
  11. I was discarding candidates over fixable flaws
  12. There are flaws transcription can’t find
  13. 70 minutes of training material disappeared in a network blink
  14. How writing “ja” as “JP” created a jargon model
  15. 4 registration paths, 0 management screens
  16. Each deployment was overwriting the other’s work
  17. Pushing unmeasured things with thresholds always fails

The insights are summarized in the notebook on mass-producing practical voices from diffusion TTS.

์›๋ฌธ์—์„œ ๊ณ„์† โ†—

์ถ”์ถœ ๋ณธ๋ฌธ ยท ์ถœ์ฒ˜: dev.to ยท https://dev.to/orca_forge/how-shosho-became-shomo-permission-character-list-was-trimming-japanese-10bi