← 피드로
[Submitted on 8 Jun 2026 (v1), last revised 30 Sep 2026 (this version, v3)]
Abstract:Work on `emergent misalignment’ shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection’ (PSM) hypothesis: during pre-training, LLMs learn to simulate different characters/perspectives, which can be elicited and refined during post-training. This paper investigates the converse phenomenon, `emergent alignment’, and uses it to support and refine the PSM and motivate a novel desideratum for alignment. We finetune a helpful-only model on broad and narrow safety tasks. To create SFT samples, we follow the `Constitutional AI’ (CAI) approach and use four constitutions which encode reasonable alignment strategies: deontology, consequentialism, virtue ethics, and aligning AIs as subordinate to human authority. For each, we show that finetuning on two narrow safety sub-categories reliably induces emergent alignment over a representative set of general safety categories, and on safety subcategories that we directly filtered-out of the data sets used for narrow alignment. To test the `PSM’ using a more fine-grained evaluation, we used a multidimensional `ethical persona’ diagnostic. For each constitutionally finetuned (broad/narrow) model, we evaluate how well their behavior matches their expected signature profile. Our results show that our CAI models acquire their expected “ethical persona” — e.g., the model narrowly fine-tuned on SFT samples created using the consequentialist constitution agrees significantly more with utilitarian than deontological beliefs. Yet our coarse and fine-grained evaluations show that there are significant differences across our (broad/narrow) finetuned CAI models in how well they project. We conclude that alignment strategies should be evaluated, not just on their (in-distribution) general safety performance, but also specifically on their degree of coarse and fine-grained projectability.
Submission history
From: Guillermo Del Pinal [view email]
[v1]
Mon, 8 Jun 2026 13:30:29 UTC (1,283 KB)
[v2]
Tue, 9 Jun 2026 13:08:40 UTC (1,283 KB)
[v3]
Wed, 30 Sep 2026 00:45:49 UTC (1,357 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.09475