더 많은 맥락, 더 큰 모델 또는 도덕적 지식? 정치 텍스트에서의 슈워츠 가치 탐지에 대한 체계적인 연구

작성자

카테고리:

← 피드로
arXiv cs.AI · V'ictor Yeste, Paolo Rosso · 2026-09-09 AI

[Submitted on 21 May 2026 (v1), last revised 7 Sep 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Detecting Schwartz values in political texts is hard: cues are often implicit, and neighboring values differ by fine distinctions. Two remedies are widely assumed to help: more surrounding document text, and explicit moral knowledge. Knowledge-based retrieval has improved benchmarks elsewhere, but whether either transfers here is untested, because published systems vary context, knowledge, and model family at once. We separate these factors under matched conditions on the ValuesML/Touché ValueEval format. The input ranges from the target sentence to a local window to the full document. Retrieval is either absent or drawn from a curated moral knowledge base, injected by early, late, or cross-attention fusion. Supervised DeBERTa-v3 encoders are compared against zero-shot LLMs from 12B to 123B. More context is not uniformly better: full-document input improves the encoders by 2.5-3.8 macro-F1 points but does not consistently help the LLMs. Retrieved knowledge helps more reliably, improving every model family and context under early fusion. A control substituting random knowledge-base entries shows the families gain differently: LLMs from relevance, encoders from exposure to the value ontology. Neither larger encoders nor larger LLMs guarantee gains, and early fusion outperforms both trainable variants. Value-sensitive NLP should evaluate context, knowledge, and model family jointly.

Submission history

From: Víctor Yeste [view email]
[v1] Thu, 21 May 2026 15:46:54 UTC (145 KB)
[v2] Fri, 22 May 2026 07:17:37 UTC (145 KB)
[v3] Thu, 11 Jun 2026 09:38:56 UTC (145 KB)
[v4] Mon, 7 Sep 2026 11:51:40 UTC (153 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.22641