← 피드로
[Submitted on 25 Sep 2026 (v1), last revised 29 Sep 2026 (this version, v2)]
Abstract:Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Submission history
From: Junxuan Li [view email]
[v1]
Fri, 25 Sep 2026 18:04:49 UTC (114 KB)
[v2]
Tue, 29 Sep 2026 01:50:57 UTC (114 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.31857