← 피드로
[Submitted on 9 Jul 2026 (v1), last revised 20 Jul 2026 (this version, v2)]
Abstract:National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weights create a presumption of trustworthiness. Such an instrument, built by and for a language community, looks like the natural choice for measuring what that community says and values. Whether such a model validly measures anything is untested at release. The evaluation of LLMs as measurement instruments is typically task-specific and stops at agreement with human coders. Agreement cannot distinguish an LLM instrument that measures a construct from one that reaches matching codes through surface correlates. We audit the presumption on a favourable case: AMALIA, Portugal’s publicly funded 9B model, coding the moral foundation of authority in European Portuguese. The \textit{recovery gap} operationalizes the audit: decompose the codebook into its theory-defined clauses, recombine them through the theory’s explicit rule, and measure how much of the original prompt’s performance the stated theory reproduces. In a pre-registered, out-of-sample study on a transcreated (English to European Portuguese) corpus, AMALIA agrees with trained coders within six points of open models eight to thirteen times its size. Yet, the recovery gap shows that only about half of coding performance on authority can be attributed to the theory. A larger multilingual LLM closes the recovery gap on the same corpus, suggesting the shortfall lies in the annotator model, not the corpus or its translation. Sovereignty earns operational and performance trust; epistemic trust requires calibration — and the audit method is inexpensive, and portable across models, languages and tasks.
Submission history
From: Manuel Pita PhD [view email]
[v1]
Thu, 9 Jul 2026 17:34:25 UTC (55 KB)
[v2]
Mon, 20 Jul 2026 07:11:20 UTC (53 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.08731
답글 남기기