Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs
Computer Science > Computer Vision and Pattern Recognition arXiv:2610.06977 (cs) [Submitted on 4 Oct 2026] Title: Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs Authors: Xiaojun Jia, Simeng Qin, Yiming Li, Jie Liao, Sensen Gao, Ke Ma, Yang Liu, Xiaochun Cao.
计算机科学 > 计算机视觉与模式识别 arXiv:2610.06977 (cs) [提交于 2026 年 10 月 4 日] 标题:针对闭源多模态大模型的可迁移对抗攻击的视觉不变性增强特征最优对齐。作者:Xiaojun Jia, Simeng Qin, Yiming Li, Jie Liao, Sensen Gao, Ke Ma, Yang Liu, Xiaochun Cao。
Abstract: Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS] embeddings. However, such coarse alignment insufficiently exploits patch-level visual structures, limiting transferability across heterogeneous closed-source MLLMs.
摘要:多模态大模型(MLLMs)仍然容易受到可迁移对抗样本的攻击,特别是在仅能访问开源代理模型的黑盒环境下。现有的目标迁移攻击主要利用全局图像级特征(如编码器 [CLS] 嵌入)来对齐对抗样本和目标样本。然而,这种粗略的对齐方式未能充分利用补丁级(patch-level)的视觉结构,从而限制了在异构闭源 MLLMs 之间的可迁移性。
We propose IAU-FOA, a visual-invariance-augmented feature optimal alignment attack with adaptive unbalanced transport, to improve targeted transferability against closed-source MLLMs. IAU-FOA aligns adversarial and target samples at both global and local levels: a cosine-based objective narrows their global semantic gap, while patch tokens are clustered into compact local patterns and matched through optimal transport for fine-grained feature alignment.
我们提出了 IAU-FOA,这是一种结合了自适应非平衡传输的视觉不变性增强特征最优对齐攻击,旨在提高针对闭源 MLLMs 的目标迁移能力。IAU-FOA 在全局和局部层面同时对齐对抗样本和目标样本:基于余弦的目标函数缩小了它们的全局语义差距,同时将补丁标记聚类为紧凑的局部模式,并通过最优传输进行匹配,以实现细粒度的特征对齐。
Balanced optimal transport enforces fixed marginal masses even for local clusters without reliable counterparts, potentially introducing misleading alignment gradients. We therefore introduce confidence-adaptive unbalanced transport to relax these constraints for weakly matched clusters, aiming to reduce unreliable local alignment and improve adversarial transferability.
平衡最优传输即使对于没有可靠对应项的局部簇也会强制执行固定的边际质量,这可能会引入误导性的对齐梯度。因此,我们引入了置信度自适应非平衡传输,以放宽对弱匹配簇的这些约束,旨在减少不可靠的局部对齐并提高对抗迁移性。
We further study the effect of input transformations and propose visual-invariance augmentation, which applies bidirectional pixel-intensity rescaling and per-channel white-balance adjustment to simulate exposure, contrast, illumination, and color-temperature variations. This strategy encourages adversarial perturbations to generalize across different visual encoders.
我们进一步研究了输入变换的影响,并提出了视觉不变性增强方法,该方法应用双向像素强度重缩放和逐通道白平衡调整,以模拟曝光、对比度、光照和色温的变化。这一策略鼓励对抗扰动在不同的视觉编码器之间实现泛化。
Extensive experiments on open-source and closed-source MLLMs show that IAU-FOA consistently outperforms state-of-the-art transferable attack methods. Code is available at this https URL.
在开源和闭源 MLLMs 上进行的广泛实验表明,IAU-FOA 的表现始终优于最先进的可迁移攻击方法。代码可在该 HTTPS 链接获取。