Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

작성자

카테고리:

← 피드로
arXiv cs.AI · Jasmine Brazilek, Harper Dunn · 2026-06-29 AI

[Submitted on 30 Apr 2026 (v1), last revised 12 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B’s preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer’s position explicit; the features that dilute it hold animal-welfare content but withhold stance.

Submission history

From: Jasmine Brazilek [view email]
[v1] Thu, 30 Apr 2026 23:59:56 UTC (286 KB)
[v2] Fri, 26 Jun 2026 15:22:15 UTC (654 KB)
[v3] Sun, 12 Jul 2026 16:52:38 UTC (45 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.26104

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다