AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

작성자

카테고리:

← 피드로
arXiv cs.AI · Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, Sergey Nikolenko · 2026-07-10 AI

[Submitted on 7 Jul 2026 (v1), last revised 14 Jul 2026 (this version, v2)]

View PDF

Abstract:We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit — did the task pass? — but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at this https URL.

Submission history

From: Vadim Lomshakov [view email]
[v1] Tue, 7 Jul 2026 11:27:43 UTC (263 KB)
[v2] Tue, 14 Jul 2026 09:13:17 UTC (263 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.06624

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다