Yes, Q-learning Helps Offline In-Context RL

작성자

카테고리:

← 피드로
arXiv cs.AI · Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov · 2026-08-15 AI

[Submitted on 24 Feb 2025 (v1), last revised 13 Aug 2026 (this version, v5)]

View PDF

Abstract:Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.

Submission history

From: Denis Tarasov [view email]
[v1] Mon, 24 Feb 2025 21:29:06 UTC (1,389 KB)
[v2] Tue, 15 Apr 2025 19:18:00 UTC (3,934 KB)
[v3] Mon, 19 May 2025 16:55:06 UTC (3,945 KB)
[v4] Mon, 25 May 2026 21:47:40 UTC (1,200 KB)
[v5] Thu, 13 Aug 2026 15:02:14 UTC (1,201 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2502.17666