RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation
Computer Science > Computer Vision and Pattern Recognition arXiv:2610.06932 (cs) [Submitted on 3 Oct 2026] Title: RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation Authors: Siyu Huang, Yueyong Chen, Xuejiao Li, Jun Zhou.
计算机科学 > 计算机视觉与模式识别 arXiv:2610.06932 (cs) [提交于 2026 年 10 月 3 日] 标题:RADC:用于视觉-语言测试时自适应的风险感知双重缓存,作者:Siyu Huang, Yueyong Chen, Xuejiao Li, Jun Zhou。
Abstract: Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching.
摘要:面向视觉-语言模型的基于缓存的测试时自适应(TTA)通常受到全局表示中背景偏差以及表示变化下基于熵的缓存准入不可靠性的阻碍。为了解决这些局限性,我们提出了 RADC,它通过可靠的双重缓存增强了原型学习。
RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference.
RADC 引入了一种语义前景缓存,它从 CLIP 表示中聚合类别一致的空间证据,产生能够补充全局缓存同时减轻背景干扰的前景原型。
To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates.
为了可靠地管理这两个缓存,高斯风险准入将多视图表示建模为对角高斯分布,并联合考虑类别分离和特征不确定性,以优先处理可靠的缓存候选对象。
RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.
RADC 将零样本逻辑值与互补的全局和前景缓存预测相结合,以实现稳健的推理。在跨域和分布外基准测试上的大量实验证明了其持续领先的性能。