Position: Behavioral Systems Require Behavioral Tests

작성자

카테고리:

← 피드로
arXiv cs.AI · Manuel Cherep, Nikhil Singh, Pattie Maes · 2026-08-20 AI

[Submitted on 31 May 2026]

View PDF HTML (experimental)

Abstract:Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time. Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them. This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their actions. We draw on lessons from the behavioral sciences to motivate this position, and propose a research agenda focused on developing rigorous behavioral tests. These include methods for recovering decision strategies from action sequences, constructing environments that isolate behavioral differences, and probing emergent dynamics in multi-agent systems. Taken together, these directions offer a roadmap for developing a science of AI behavior.

Submission history

From: Manuel Cherep [view email]
[v1] Sun, 31 May 2026 02:09:20 UTC (331 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.18081