Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring

작성자

카테고리:

← 피드로
arXiv cs.AI · Yongchao Wu, Aron Henriksson · 2026-07-11 AI

[Submitted on 8 Apr 2026 (v1), last revised 8 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Activation-based steering enables inference-time personalization of large language models, but its effects in educational applications are not well understood. We study activation-based persona vectors representing seven character traits in short-answer generation and automated scoring on the ASAP-SAS benchmark, across three language models spanning dense and mixture-of-experts architectures. Persona steering lowers answer quality overall, with much larger effects on open-ended English Language Arts (ELA) prompts than on factual science prompts. Interpretive and argumentative tasks are particularly sensitive, showing up to 11$\times$ larger degradation. On the scoring side, we observe predictable valence-aligned calibration shifts: “evil” and “impolite” scorers grade more harshly, while “good” and “optimistic” scorers grade more leniently. ELA tasks are 2.5-3$\times$ more susceptible to scorer personalization than science tasks, and the mixture-of-experts model shows roughly 6$\times$ larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring. Our findings highlight the need for task- and architecture-aware calibration when deploying personalized models in educational settings.

Submission history

From: Yongchao Wu Dr. [view email]
[v1] Wed, 8 Apr 2026 13:55:32 UTC (888 KB)
[v2] Wed, 8 Jul 2026 20:25:11 UTC (750 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.07102

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다