Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light

작성자

카테고리:

← 피드로
arXiv cs.AI · Mani Hamidi, Terrence W. Deacon · 2026-06-23 AI

[Submitted on 15 Jul 2025 (v1), last revised 21 Jun 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the question of what constitutes agency. Reinforcement learning (RL) offers arguably the most explicit formulation of agent-environment interaction, built on three core tenets: the environment as a Markov decision process, learning as policy optimization, and the agent as a maximizer of scalar reward. Recent work has called to revise these tenets: reconceptualizing learning as adaptation rather than optimization, broadening goals beyond scalar reward, and noting the absence of a formal theory of the agent in a formalism that so heavily emphasizes the environment. We argue that the artificial life community is uniquely positioned to illuminate this critique and concretize an alternative. We draw on open-ended novelty search as a complementary model of adaptation and goal-directed behavior beyond reward optimization, and ground such evolutionary dynamics in thermodynamic theories of origin-of-life and agency, toward a more biologically faithful and formally grounded account of what it is to be an adaptive agent.

Submission history

From: Mani Hamidi [view email]
[v1] Tue, 15 Jul 2025 16:53:14 UTC (94 KB)
[v2] Fri, 18 Jul 2025 05:07:38 UTC (94 KB)
[v3] Mon, 28 Jul 2025 18:54:04 UTC (94 KB)
[v4] Sun, 21 Jun 2026 09:03:48 UTC (126 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2507.11482

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다