SciLitBench: LLM 기반 체계적 문헌 검토를 위한 벤치마크 및 설계 원칙

작성자

카테고리:

← 피드로
arXiv cs.AI · Miguel Zabaleta, Baihan Lin · 2026-09-09 AI

[Submitted on 29 Aug 2026]

View PDF HTML (experimental)

Abstract:Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30\% of annotated evaluation evidence and 25\% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.

Submission history

From: Miguel Zabaleta [view email]
[v1] Sat, 29 Aug 2026 00:04:34 UTC (3,833 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.05505