Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

작성자

카테고리:

← 피드로
arXiv cs.AI · Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant · 2026-06-09 AI

[Submitted on 5 Jun 2026]

View PDF HTML (experimental)

Abstract:LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their susceptibility to relying on in context-information, and their steerability to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of many generalist LLMs and safety-specific judges, and investigate the impact of task demonstrations, novel in-context information, and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to adjust their evaluations if the context or safety definition contradicts their prior.

Submission history

From: Anissa Alloula [view email]
[v1] Fri, 5 Jun 2026 22:11:26 UTC (2,547 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.07874

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다