Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

작성자

카테고리:

← 피드로
arXiv cs.AI · Izumi Takahara, Teruyasu Mizoguchi · 2026-07-13 AI

[Submitted on 10 Jul 2026]

View PDF HTML (experimental)

Abstract:Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis–test–evidence–belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.

Submission history

From: Izumi Takahara [view email]
[v1] Fri, 10 Jul 2026 08:39:30 UTC (1,148 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.09195

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다