HyQuant: Hybrid-Precision Quantization for LLM Attention

작성자

카테고리:

← 피드로
arXiv cs.AI · Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang · 2026-09-14 AI

[Submitted on 28 Aug 2026 (v1), last revised 16 Sep 2026 (this version, v3)]

Authors:Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang

View PDF HTML (experimental)

Abstract:Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: this https URL .

Submission history

From: Jiatong Ding [view email]
[v1] Fri, 28 Aug 2026 03:30:28 UTC (16,996 KB)
[v2] Fri, 11 Sep 2026 03:12:02 UTC (16,996 KB)
[v3] Wed, 16 Sep 2026 01:12:44 UTC (16,996 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.27875