“미국 영어”가 어떻게 기본값이 되는가? LLM 교육 과정 전반에 걸친 미국 영어에 대한 구조적 편향 분석

작성자

카테고리:

← 피드로
arXiv cs.AI · Mir Tafseer Nayeem, Davood Rafiei · 2026-09-29 AI

[Submitted on 5 Apr 2026 (v1), last revised 28 Sep 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose “English (US)” as a primary English setting despite the global diversity of English. We ask: How does “English (US)” become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)–British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure –> representation –> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.

Submission history

From: Mir Tafseer Nayeem [view email]
[v1] Sun, 5 Apr 2026 17:59:34 UTC (2,505 KB)
[v2] Mon, 28 Sep 2026 00:34:52 UTC (3,097 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.04204