Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

작성자

카테고리:

← 피드로
arXiv cs.AI · Yongjin Yang, Jiarui Liu, Yinghui He, Lezhen Zhang, Bernhard Sch"olkopf, Zhijing Jin · 2026-06-25 AI

[Submitted on 23 Jun 2026 (v1), last revised 27 Jun 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science. However, the training curriculum (how often each domain is sampled) is typically fixed or hand-tuned, even though reasoning skills transfer unevenly across domains. Existing learnability-based curricula adapt to where the policy is currently improving, but are blind to whether a gradient step on the selected domain benefits the remaining domains. In this paper, we propose Transfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose updates broadly benefit the rest of the training suite. TAC repurposes signals already produced by RL training: per-domain advantages capture local learnability, and projected gradients, taken from the GRPO step being computed, estimate cross-domain transferability via gradient-geometry alignment, at negligible cost (<1% wall-clock overhead). Across a six-domain reasoning suite, TAC achieves the best macro-averaged accuracy on both Qwen3-1.7B and Llama3.2-3B, outperforming proportional random sampling, a hand-designed schedule, and a learnability-only bandit, and improving over the last of these by up to 2.8 points (10% relative). Ablations show performance degrades sharply when the transferability term is removed, and TAC remains robust on imbalanced training mixtures where learnability-only curricula over-commit to dominant domains. Our findings establish cross-domain transferability as a key signal for curriculum design in multi-domain RLVR.

Submission history

From: Yongjin Yang [view email]
[v1] Tue, 23 Jun 2026 21:10:29 UTC (670 KB)
[v2] Sat, 27 Jun 2026 02:34:55 UTC (670 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.25178

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다