시코팬시가 아닌 수용성: 언어 모델의 정의와 참여의 구별

작성자

카테고리:

← 피드로
arXiv cs.AI · Calvin Isley, Johann Gaebler, Max Lamparth, Julia Minson, Sharad Goel · 2026-09-23 AI

[Submitted on 22 Sep 2026]

View PDF HTML (experimental)

Abstract:A central concern with language models is sycophancy: their tendency to defer to users’ views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy are also characteristic of conversational receptiveness, a construct from social psychology shown to improve interactions across disagreement. We argue that this overlap creates a construct-validity problem for social sycophancy evaluations. Using a popular moral-advice dataset, we find that responses classified as more socially sycophantic are also more receptive. Further, increasing the receptiveness of human-written responses—while preserving their substantive conclusions—causes them to be classified as more socially sycophantic. This tight coupling raises the possibility that social sycophancy evaluations inadvertently penalize desirable behavior. In a preregistered experiment comparing substantively equivalent responses, participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors. The same overall pattern persists even among participants who believe the original question asker is in the wrong. Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.

Submission history

From: Calvin Isley [view email]
[v1] Tue, 22 Sep 2026 15:27:16 UTC (93 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.26579