BRIDGE: Predicting Human Task Completion Time From Model Performance

작성자

카테고리:

← 피드로
arXiv cs.AI · Fengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau, Siva Reddy, Hugo Larochelle · 2026-07-03 AI

[Submitted on 6 Feb 2026 (v1), last revised 2 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on direct human task completion time annotations are costly, noisy, and difficult to scale across benchmarks. In this work, we propose BRIDGE, a unified psychometric framework that learns a latent difficulty scale from model responses and anchors it to human task completion time. Using a two-parameter logistic Item Response Theory model, we jointly estimate latent task difficulty and model capability from model performance data across multiple benchmarks. We demonstrate that latent task difficulty varies linearly with the logarithm of human completion time, allowing human task completion time to be inferred for new benchmarks from model performance alone. Leveraging this alignment, we forecast frontier model capabilities in terms of human task length and independently reproduce METR’s exponential scaling results, with the 50% solvable task horizon doubling approximately every 6 months.

Submission history

From: Fengyuan Liu [view email]
[v1] Fri, 6 Feb 2026 23:36:11 UTC (1,551 KB)
[v2] Thu, 2 Jul 2026 11:08:50 UTC (1,555 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2602.07267

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다