← 피드로
[Submitted on 29 Jul 2026]
Abstract:Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher’s defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward’s feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher’s model, emitting structured preference constraints (understanding statements) that keep the teacher’s model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher’s error about the learner is concentrated in a particular direction rather than spread evenly.
Submission history
From: Jack Mirenzi [view email]
[v1]
Wed, 29 Jul 2026 18:44:01 UTC (925 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.11229
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.