희소 중첩에서 LLM 판사 검증: 추론에서 설계까지

작성자

카테고리:

← 피드로
arXiv cs.AI · Junxuan Li, Arko Mukherjee, Soumyabrata Pal · 2026-09-29 AI

[Submitted on 25 Sep 2026 (v1), last revised 29 Sep 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

Submission history

From: Junxuan Li [view email]
[v1] Fri, 25 Sep 2026 18:04:49 UTC (114 KB)
[v2] Tue, 29 Sep 2026 01:50:57 UTC (114 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.31857