PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

PerceptionBench:评估多模态大语言模型中的原子级视觉感知能力

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). 我们推出了 PerceptionBench,这是一个专门为评估多模态大语言模型(MLLMs)的原子级视觉感知能力而设计的基准测试。

Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. 现有的基准测试往往无法将感知能力独立出来:整体评估将感知错误与推理或领域知识的缺失混为一谈,而应用驱动的基准测试仅涵盖由启发式设计所形成的狭窄且碎片化的领域。

To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. 为了解决这些局限性,PerceptionBench 采用了自下而上的方法:通过诊断前沿 MLLMs 在 42 个现有基准测试中响应的最早失败点,我们构建了一个错误分类法,其感知分支定义了十种原子级感知能力。

Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. 在该分类法的指导下,我们构建了 3,000 个经过验证的问题,这些问题答案简短且无歧义,每个问题都针对单一能力进行隔离,其难度源于感知而非推理或知识。

Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. 对 16 个前沿 MLLMs 的基准测试结果显示,原子级感知问题在很大程度上仍未得到解决——没有模型达到 60% 的准确率,与感知相关的幻觉是平均表现最弱的能力,且相似的总分掩盖了截然不同的能力分布特征。

PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs. 因此,PerceptionBench 为衡量和诊断 MLLMs 的视觉感知边界提供了一个能力层面的标准。