Redistribution-based Cost Inference Improves Sparse Safe Offline RL

작성자

카테고리:

← 피드로
arXiv cs.AI · Ebenezer Gelo (University of the Witwatersrand), Geraud Nangue Tasse (University of the Witwatersrand), Steven James (University of the Witwatersrand), Benjamin Rosman (University of the Witwatersrand) · 2026-08-13 AI

[Submitted on 12 Aug 2026]

View PDF HTML (experimental)

Abstract:Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset. We show that return-equivalent redistribution preserves the feasible policy set and the optimal Lagrangian in a CMDP, establishing that the transformation is lossless in theory while yielding better-conditioned cost critic learning in practice. Experiments on highway driving and robotic manipulation demonstrate substantially lower violation rates than sparse and classifier-based baselines, with robustness to heterogeneous dataset compositions and label noise.

Submission history

From: Ebenezer Gelo [view email]
[v1] Wed, 12 Aug 2026 17:53:15 UTC (7,546 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.12306

코멘트

답글 남기기