GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding

작성자

카테고리:

← 피드로
arXiv cs.AI · Fanxu Meng · 2026-07-22 AI

[Submitted on 14 May 2026 (v1), last revised 21 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matches the H100 roofline almost perfectly. Its trained weights, however, expose only one decoding path – an absorbed MQA form – which ties efficient inference to H100-class compute-bandwidth ratios, forfeits tensor parallelism along the head axis, and yields no Multi-Token Prediction (MTP) gain on commodity inference GPUs such as the export-restricted H20. We propose Group-Query Latent Attention (GQLA), a minimal modification of MLA whose trained weights expose two algebraically equivalent decoding paths over the same parameters: an MQA-absorb path identical to MLA’s, and a GQA path with a per-group expanded cache. The runtime picks the path that matches the target hardware – no retraining, no custom kernels – so a single set of GQLA weights pins the rooflines of both H100 (MQA-absorb, s_q=1) and H20 (GQA + MTP, s_q=2), while supporting up to 8-way zero-redundancy tensor parallelism on the GQA path. To avoid pretraining from scratch we extend TransMLA into TransGQLA, which converts a pretrained GQA checkpoint into a GQLA model; on LLaMA-3-8B it compresses the per-token KV cache to 28.125% of the GQA baseline on the MQA-absorb path while structurally preserving GQA-level traffic on the per-group path.

Submission history

From: Meng Fanxu [view email]
[v1] Thu, 14 May 2026 15:50:01 UTC (665 KB)
[v2] Wed, 27 May 2026 07:33:41 UTC (663 KB)
[v3] Tue, 21 Jul 2026 12:19:20 UTC (684 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.15250

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다