HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

작성자

카테고리:

← 피드로
arXiv cs.AI · Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar · 2026-09-28 AI

[Submitted on 24 Sep 2026]

View PDF HTML (experimental)

Abstract:Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.

Submission history

From: Aditya Kumaran [view email]
[v1] Thu, 24 Sep 2026 21:33:11 UTC (830 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.30571