Surgical Video Generation From Diffusion to World Models: A Survey
Surgical Video Generation From Diffusion to World Models: A Survey
手术视频生成:从扩散模型到世界模型综述
Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance.
摘要: 手术视频数据是术中感知、手术流程理解和机器人决策模型的主要训练资源。然而,临床数据的获取仍受限于隐私、成本和类别不平衡等问题。
Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework.
手术视频生成已成为解决数据稀缺问题的变革性方法,并为手术模拟、培训和机器人策略学习奠定了基础。该领域发展迅速,但目前尚缺乏清晰的概念框架。
This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes.
本综述将 2024 年至 2026 年的文献归纳为三类:无条件生成、条件生成和世界模型生成,揭示了该任务定义的根本性转变——即从合成视觉上逼真的帧图像,转向对手术场景因果动力学的建模。
We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field.
我们审视了像素级保真度与临床合理性之间持续存在的差距,并确定了泛化性、物理真实性、可控性和可解释性等瓶颈。我们进一步总结了代表性方法在公开数据集上的实验结果,为该领域提供定量参考。
This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.
本综述对手术视频生成领域的现状和面临的挑战进行了结构化概述,为从事智能感知、多模态融合、生成式 AI 和手术数据科学交叉领域研究的学者提供了参考。