LLM Scheming Inversely Scales with Pretraining Language Coverage

작성자

카테고리:

← 피드로
arXiv cs.AI · Nathan Truong, Aryan Panda, Rayming Ye, Zoe Sun, Maheep Chaudhary · 2026-07-29 AI

[Submitted on 9 Jun 2026]

View PDF HTML (experimental)

Abstract:With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming — the covert pursuit of misaligned objectives while feigning alignment — in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2\% higher scores compared to high-resource languages on a five-category scheming index. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.

Submission history

From: Maheep Chaudhary [view email]
[v1] Tue, 9 Jun 2026 06:02:50 UTC (1,026 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.24769

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다