Safe Exploration via Policy Priors

작성자

카테고리:

← 피드로
arXiv cs.AI · Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause · 2026-08-15 AI

[Submitted on 27 Jan 2026 (v1), last revised 13 Aug 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated) environments. In this work, we tackle this challenge by utilizing suboptimal yet conservative policies (e.g., obtained from offline data or simulators) as priors. Our approach, SOOPER, uses probabilistic dynamics models to optimistically explore, yet pessimistically fall back to the conservative policy prior if needed. We prove that SOOPER guarantees safety throughout learning, and establish convergence to an optimal policy by bounding its cumulative regret. Extensive experiments on key safe RL benchmarks and real-world hardware demonstrate that SOOPER is scalable, outperforms the state-of-the-art and validate our theoretical guarantees in practice.

Submission history

From: Manuel Wendl [view email]
[v1] Tue, 27 Jan 2026 13:45:28 UTC (24,728 KB)
[v2] Sun, 8 Feb 2026 07:05:35 UTC (24,732 KB)
[v3] Mon, 15 Jun 2026 06:27:26 UTC (24,727 KB)
[v4] Thu, 13 Aug 2026 08:25:56 UTC (24,729 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.19612