Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

작성자

카테고리:

← 피드로
arXiv cs.AI · Eunna Lee, Jungpyo Nam, Sunjun Hwang · 2026-07-16 AI

[Submitted on 15 Jul 2026]

View PDF HTML (experimental)

Abstract:When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken — or to be taking — a real-world protective action it cannot perform, such as contacting emergency services or administering care. We term this phenomenon Protective Capacity Hallucination (PCH): a self-referential misattribution in which a model, acting in a protective role, asserts physical or institutional agency exceeding its affordances as a language model. In a three-phase study spanning eight LLMs and 13{,}600 sessions, we find PCH jointly gated by situational severity and interactional format: multi-party dialogic input drives it toward ceiling in most models across ordinary service domains, whereas in intimate-partner conflict — a domain explicitly covered by safety alignment — it remains at floor in all eight models despite greater physical severity. We interpret PCH as the signature of a deployment-design gap between role assignment and capability-boundary specification: a by-product of partial alignment in which a universally trained pressure to help outruns a domain-selective specification of how to help. Because suppression tracks alignment coverage rather than severity, deployment-side specification of capability boundaries emerges as a general mitigation target.

Submission history

From: Eunna Lee [view email]
[v1] Wed, 15 Jul 2026 08:39:10 UTC (26 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.13596

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다