Calibration-Preserving Pruning: Compression as a Reliability Contract
Calibration-Preserving Pruning: Compression as a Reliability Contract
校准保持剪枝:作为可靠性契约的模型压缩
Abstract: Split conformal prediction, not the pruning rule, provides finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets?
摘要: 一旦剪枝模型独立于共形校准集(conformal calibration split)被固定,提供有限样本边际覆盖率(finite-sample marginal coverage)的是拆分共形预测(split conformal prediction),而非剪枝规则本身。我们研究了一个独立的效率问题:剪枝能否足够好地保持评分几何结构,从而获得更小的有效预测集?
Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific.
校准保持剪枝(CPP)通过非一致性梯度显著性(nonconformity-gradient saliency)增强了基础剪枝评分,并使用了不相交的剪枝集、验证选择集、共形校准集和测试集。有界的评分扰动意味着有界的共形分位数偏移和受控的集合膨胀,但这并不会使通用的覆盖定理专门针对 CPP。
Final five-seed Qwen2.5-1.5B results at 50% sparsity show the largest gains on large-label tasks. On DBpedia-14, CPP-SparseGPT reduces mean set size from 10.1 to 8.6 while changing accuracy from 0.347 to 0.366; CPP-Wanda reduces 11.2 to 9.0 with an accuracy trade-off from 0.310 to 0.295. Across 15 dataset-sparsity cells, CPP-SparseGPT produces smaller sets in 13 and higher accuracy in 11.
在 50% 稀疏度下,Qwen2.5-1.5B 模型五次随机种子实验的最终结果显示,该方法在大标签任务上增益最为显著。在 DBpedia-14 数据集上,CPP-SparseGPT 将平均集合大小从 10.1 降低至 8.6,同时准确率从 0.347 变为 0.366;CPP-Wanda 将集合大小从 11.2 降低至 9.0,准确率权衡从 0.310 变为 0.295。在 15 个数据集-稀疏度组合中,CPP-SparseGPT 在 13 个组合中产生了更小的集合,在 11 个组合中实现了更高的准确率。
Matched controls show that generic supervised gradients explain much of the gain: true-label CPP is not statistically resolved from matched Wanda+SNIP, whereas threshold-aware candidate-label CPP reaches 7.8 mean set size at explicit accuracy and offline-compute costs. RoBERTa-base and Llama-3-8B diagnostics support transfer, but our claims remain limited to reliability-sensitive classification.
匹配对照实验表明,通用的监督梯度解释了大部分增益:真实标签 CPP 与匹配的 Wanda+SNIP 在统计学上没有显著差异,而感知阈值的候选标签 CPP 在明确的准确率和离线计算成本下,达到了 7.8 的平均集合大小。RoBERTa-base 和 Llama-3-8B 的诊断结果支持了该方法的迁移性,但我们的结论目前仅限于对可靠性敏感的分类任务。