AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

Computer Science > Machine Learning arXiv:2610.06927 (cs) [Submitted on 3 Oct 2026] Title:AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD Authors:Sara Abdali, Jongwoo Ko, Pashmina Cameron View a PDF of the paper titled AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD, by Sara Abdali and 2 other authors View PDF HTML (experimental)

计算机科学 > 机器学习 arXiv:2610.06927 (cs) [提交于 2026 年 10 月 3 日] 标题:AttSVD:通过注意力引导 SVD 实现提示自适应低秩 KV 缓存压缩 作者:Sara Abdali, Jongwoo Ko, Pashmina Cameron 查看论文 PDF,标题为《AttSVD:通过注意力引导 SVD 实现提示自适应低秩 KV 缓存压缩》,作者 Sara Abdali 及其他 2 位作者 查看 PDF HTML(实验性)

Abstract:The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the “feature” axis. We therefore propose AttSVD, a new “interpretable” low-rank compression whose basis is derived from each prompt’s own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank.

摘要:自回归 Transformer 的键值 (KV) 缓存随上下文长度线性增长,并在长上下文中占据内存主导地位。大多数无需训练的补救措施会剔除低重要性标记,这在序列轴上是一种不可逆的选择。我们转而保留每个标记,并沿“特征”轴以更廉价的方式存储它。因此,我们提出了 AttSVD,这是一种新的“可解释”低秩压缩方法,其基底源自每个提示自身的注意力几何结构:一种在线的、针对每个提示的截断 SVD,仅保留注意力实际读取的方向,从而按保留秩的比例减少持久的每头 KV 内存。

We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes.

我们针对短生成和长生成场景提出了两种解码时缓存策略:累积式和流式。此外,我们提出了两种使压缩具有自适应性的改进方案。一种基于矩阵的能量规则独立地调整 Logit 空间和注意力权重的大小。一种注意力感知基底仅在注意力实际读取的空间中进行截断,从而同时保留了注意力 Logit 和注意力输出。同样的因素也提供了关于有效秩和注意力所消耗几何结构的免费、每头可解释性见解。

Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.06927 [cs.LG] (or arXiv:2610.06927v1 [cs.LG] for this version)

在多个模型上,无论是在智能体基准测试还是完整的 LongBench 套件中,AttSVD 在使用高达 50% 的 KV 缓存内存的同时,保持了与密集缓存相当的性能。学科:机器学习 (cs.LG);人工智能 (cs.AI) 引用方式:arXiv:2610.06927 [cs.LG](或此版本的 arXiv:2610.06927v1 [cs.LG])