Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

Obshazard-bench:用于实时灾害情报的原始地球观测流多模态基础模型基准测试


Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenarios where hazards evolve rapidly and decisions must be made under strict time constraints.

多模态大语言模型(MLLMs)正越来越多地被用于解读地球观测数据,但其支持现实世界灾害应急响应的能力仍未得到充分评估。现有的遥感基准测试大多依赖于静态的、事后的、经专家处理的产品(如网格化再分析数据),这些数据难以与灾害快速演变且必须在严格时间限制下做出决策的实际灾害场景相匹配。

To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs. Unlike image-centric or post-event benchmarks, Obshazard-bench directly integrates raw, high-frequency satellite sounding streams from diverse satellite sensors with concurrent ground-station observations, historical disaster records, and socio-economic indicators, bypassing delayed expert-processing and physical-inversion pipelines.

为了弥补这一差距,我们推出了 Obshazard-bench,这是一个用于评估 MLLM 灾害情报的实时、观测驱动型基准测试。与以图像为中心或事后评估的基准不同,Obshazard-bench 直接整合了来自多种卫星传感器的原始高频卫星探测流,以及同步的地面站观测数据、历史灾害记录和社会经济指标,从而绕过了延迟的专家处理和物理反演流程。

The benchmark covers 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historically documented extreme-event cases and thousands of lifecycle-oriented VQA samples. Moreover, Obshazard-bench further defines a three-stage evaluation taxonomy aligned with the operational disaster workflow: Predictive Crisis Anticipation for pre-disaster risk detection and early forecasting, Active Evolution Reasoning for in-situ disaster tracking and termination prediction, and Multi-faceted Impact Quantification for post-disaster magnitude deduction, humanitarian burden estimation, and socio-economic impact assessment.

该基准测试涵盖了全球 60 多个国家的 8 大类灾害和 28 个子类别,纳入了超过 120 个有记录的历史极端事件案例和数千个面向生命周期的视觉问答(VQA)样本。此外,Obshazard-bench 还定义了一个与灾害业务流程相一致的三阶段评估分类体系:用于灾前风险检测和早期预报的“预测性危机预判”、用于现场灾害跟踪和终止预测的“主动演变推理”,以及用于灾后规模推断、人道主义负担估算和社会经济影响评估的“多维度影响量化”。

Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning.

在代表性的通用基础模型和地球科学专用基础模型上进行的实验表明,这些模型在将原始多通道物理观测转化为具有时间基础和决策相关性的灾害推理方面存在显著局限性。