Agent Safety Is Action Alignment

작성자

카테고리:

← 피드로
arXiv cs.AI · Shawn Li, Yue Zhao · 2026-06-30 AI

[Submitted on 27 Jun 2026]

View PDF HTML (experimental)

Abstract:Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user’s behalf. To keep them safe, practitioners imported the chatbot-era recipe (train the model to refuse unsafe inputs) into the agentic setting, and treat the resulting capability loss as a manageable “alignment tax.” We argue this is a \emph{category error}. Refusal is a primitive for \emph{content safety}, where the harm is in the model’s output and is therefore a learnable function of it. Agentic harm is different in kind: it lies not in any output but in the relation between the authority an action exercises and the authority the user granted, which is absent from the text the model sees. Importing content-safety methods into this regime does not trade capability for safety; it pays capability and buys negative security. We support this with three lines of evidence spanning the autonomy spectrum: defense-trained models learn surface patterns rather than intent; the same training collapses multi-step agents before any threat appears while leaving them exploitable; and even undefended frontier models exceed granted authority under ordinary use. We conclude that action safety cannot be installed in weights. It must be expressed as \emph{least privilege}, enforced \emph{outside} the model at the action boundary, and evaluated as \emph{action alignment} (a relational, deployment-conditioned property) rather than a refusal score.

Submission history

From: Li Li [view email]
[v1] Sat, 27 Jun 2026 05:26:43 UTC (142 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.28739

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다