Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

작성자

카테고리:

← 피드로
arXiv cs.AI · Chris Han, Pengzhi Gao, Pei Fu, Jian Luan · 2026-08-13 AI

[Submitted on 11 Aug 2026 (v1), last revised 12 Aug 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5. We further investigate on-policy distillation (OPD) and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation. We release the models and code to facilitate future research.

Submission history

From: Chris Han [view email]
[v1] Tue, 11 Aug 2026 11:30:38 UTC (1,345 KB)
[v2] Wed, 12 Aug 2026 10:06:17 UTC (1,345 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.10812

코멘트

답글 남기기