← 피드로
[Submitted on 14 May 2026]
Abstract:Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment. Recent advances in policy expressiveness and scale have intensified this challenge, leading to a rapidly growing but conceptually fragmented body of work on RL policy verification. This survey provides a unifying perspective on RL verification methods. We introduce a taxonomy that clarifies relationships among existing approaches along three axes: verification paradigm (formal versus probabilistic), temporal scope (step-wise versus multi-step), and guarantees strength. Beyond taxonomy, we unify underlying theoretical foundations, make implicit assumptions and limitations explicit, and identify emerging directions.
Submission history
From: Luca Marzari [view email]
[v1]
Thu, 14 May 2026 12:32:57 UTC (2,356 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.16210
답글 남기기