Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

작성자

카테고리:

← 피드로
arXiv cs.AI · Benjamin Poole, Minwoo Lee · 2026-07-11 AI

[Submitted on 8 Jul 2026]

View PDF

Abstract:Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline training in fully sequential decision-making environments. We propose Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies. We adapt Safety Gymnasium environments to be a principled testbed for alignment evaluation, demonstrating improved aptitude and up to a 98% reduction in misalignment across a range of imitation learning algorithms. FMR remains robust in limited data regimes, even when learning from scarce aligned and uninformative noisy demonstrations.

Submission history

From: Benjamin Poole [view email]
[v1] Wed, 8 Jul 2026 18:47:29 UTC (7,038 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.07859

코멘트

답글 남기기