Explorerโ€บData Scienceโ€บMachine Learning
Research PaperResearchia:202609.30063

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Jiale Chen

Abstract

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and ...

Submitted: September 30, 2026Subjects: Machine Learning; Data Science

Description / Details

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.


Source: arXiv:2609.38121v1 - http://arxiv.org/abs/2609.38121v1 PDF: https://arxiv.org/pdf/2609.38121v1 Original Link: http://arxiv.org/abs/2609.38121v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 30, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms | Researchia