HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
HeadGuard:针对低比特视觉语言模型(VLM)KV 缓存量化的选择性头部保护技术
Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries.
摘要: 低比特键值(KV)缓存量化虽然节省了存储空间,但往往会严重降低视觉语言模型(VLM)的准确性。我们引入了 HeadGuard,这是一种可组合的头部保护方法,通过固定的高精度掩码来增强基础 KV 缓存量化器。该方法通过离线计算图像敏感度和输出敏感度分数来筛选物理 KV 头部,在主要实验中约有 1/8 的头部受到保护;这些头部的图像键(keys)以及可选的数值(values)保持为 bfloat16 (BF16) 格式,而基础量化器则负责处理未受保护的图像条目。
Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits, the six-task discriminative mean over eight models rises from 0.436 to 0.580 on the weakest base; protection can also improve generated answers and caption fidelity to BF16 outputs. Mean accuracy gains persist across all three quantizers with both tested calibration datasets.
在涵盖 8 个 VLM 模型、3 种基础量化器和 8 个基准测试(6 个判别式和 2 个生成式)的实验中,HeadGuard 在较弱的量化器上恢复了大部分丢失的准确率,其中在 Qwen 和 InternVL 模型上的提升最为显著。在 2 比特量化下,8 个模型在 6 项判别任务中的平均准确率从最弱基础量化器的 0.436 提升至 0.580;此外,该保护机制还能改善生成答案的质量以及字幕与 BF16 输出的保真度。在所有三种量化器和两种测试校准数据集下,平均准确率的提升均保持稳定。
Keys-only protection retains substantial recovery at lower modeled storage cost. Evaluated through simulated quantization, HeadGuard offers a composable way to improve low-bit VLM accuracy without replacing the underlying量化器.
仅保护键(Keys-only)的策略在降低存储成本的同时,依然保留了显著的准确率恢复效果。通过模拟量化评估,HeadGuard 提供了一种无需替换底层量化器即可提升低比特 VLM 准确性的可组合方案。