Uncertainty-Aware Decision Making in Multimodal Large Language Models

Uncertainty-Aware Decision Making in Multimodal Large Language Models

多模态大语言模型中具备不确定性意识的决策制定

Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a question that is not answerable from the supplied evidence.

摘要: 多模态大语言模型(MLLMs)越来越多地回答那些正确性依赖于视觉、文本、时间、音频、文档、图表或具身证据的问题。因此,它们的失败不仅仅局限于语言层面。一个流畅的回答可能掩盖了输入质量差、感知错误、基础不牢、模态间冲突、推理不稳定、分布偏移,或者问题本身无法从所提供的证据中得到解答等问题。

This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action. We review work on token and logit uncertainty, semantic disagreement, perturbation instability, grounding and attribution scores, verbalized confidence, verifier and judge scores, conformal prediction, selective answering, abstention, clarification, retrieval, self-checking, and escalation.

本综述围绕一个以决策为中心的框架组织了关于具备不确定性意识的 MLLM 的文献:不确定性来源产生可观测的信号,信号必须经过校准或风险控制,而校准后的不确定性应决定系统的行动。我们回顾了关于 Token 和 Logit 不确定性、语义分歧、扰动不稳定性、基础与归因评分、口头置信度、验证器与判别器评分、共形预测、选择性回答、弃权、澄清、检索、自我检查以及升级处理等方面的研究工作。

The central argument is that uncertainty should not be evaluated only as a confidence number; it should be evaluated by whether it improves behavior under insufficient, conflicting, shifted, or high-risk multimodal evidence. We position this survey against text-only uncertainty and abstention surveys, broad MLLM surveys, MLLM hallucination surveys, and safety-oriented reviews.

本文的核心观点是,不确定性不应仅仅被评估为一个置信度数值;它应该通过其是否能在证据不足、冲突、分布偏移或高风险的多模态证据下改善系统行为来进行评估。我们将本综述与纯文本不确定性和弃权综述、广泛的 MLLM 综述、MLLM 幻觉综述以及面向安全的综述进行了对比定位。

We conclude with open problems in source-aware decomposition, action-aware benchmarks, calibration under shift, black-box uncertainty estimation, broader modality coverage, reproducible reporting, and human-centered uncertainty communication.

最后,我们总结了在源感知分解、行动感知基准测试、偏移下的校准、黑盒不确定性估计、更广泛的模态覆盖、可重复的报告以及以人为中心的不确定性沟通等领域中尚待解决的问题。