Benchmarking LLM Competence on Logical Inference over Probability Operators

작성자

카테고리:

← 피드로
arXiv cs.AI · Nayera Hasan, Jack Greff, Alvin Grissom II · 2026-08-05 AI

[Submitted on 29 Jul 2026 (v1), last revised 3 Aug 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators–inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model’s accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.

Submission history

From: Alvin Grissom II [view email]
[v1] Wed, 29 Jul 2026 19:22:09 UTC (231 KB)
[v2] Fri, 31 Jul 2026 02:03:33 UTC (231 KB)
[v3] Mon, 3 Aug 2026 19:14:45 UTC (231 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.27405

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다