KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

작성자

카테고리:

← 피드로
arXiv cs.AI · Yijia Fang, Yiqing Feng, Bingyu Li, Mingxun Zhou · 2026-10-01 AI

[Submitted on 28 May 2026 (v1), last revised 30 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Relay and reseller APIs mediate access to large language models (LLMs), but users cannot directly verify which model serves them. We introduce \name, a black-box auditing protocol based on stable factual recall near the knowledge boundary, including repeatable wrong answers. KBF generates benign, renewable probes and calibrates audit decisions against reference self-variation. Across 16 production endpoints, KBF detects all 155 economically relevant substitutions without rejecting any of the 16 same-reference controls. KBF remains robust to deployment variation and reaches 95\% TPR at a substitution rate as low as 15\% in mixed-routing simulations. Field audits flag 7 of 28 endpoints across six platforms as statistically inconsistent with their references. After reference enrollment, even GPT-6 Astra costs only approximately \$0.67 per online audit at the recorded API prices.

Submission history

From: Yijia Fang [view email]
[v1] Thu, 28 May 2026 07:40:24 UTC (208 KB)
[v2] Sat, 25 Jul 2026 16:02:05 UTC (215 KB)
[v3] Wed, 30 Sep 2026 14:02:53 UTC (577 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.29524