Surgical Video Generation: From Diffusion to World Models — A Survey

Surgical Video Generation: From Diffusion to World Models — A Survey


Authors: Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang


Submitted: 26 August 2026 (arXiv:2608.26214 [cs.CV])


Accepted for oral presentation at: 2026 3rd International Conference on Intelligent Perception and Pattern Recognition (IPPR 2026)


Abstract


Surgical video data constitutes the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition is constrained by privacy concerns, high collection costs, and class imbalance. Surgical video generation has emerged as a transformative approach to address data scarcity, and it now serves as a foundation for surgical simulation, training, and robotic policy learning. Despite rapid developments, the field lacks a clear conceptual framework. This survey organizes the 2024–2026 literature into three categories: unconditional generation, conditional generation, and world model generation. In doing so, it reveals a fundamental shift in task definition—from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as critical bottlenecks. Furthermore, we summarize experimental results of representative methods on public datasets to provide a quantitative reference. This survey offers a structured overview of the current state and open challenges, serving as a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.


1. Introduction


The acquisition of high-quality surgical video data remains one of the most significant bottlenecks in developing intelligent systems for operating rooms. Unlike general-purpose video datasets, surgical recordings are inherently sensitive, expensive to obtain, and often exhibit severe class imbalance—with rare but critical events being underrepresented. These constraints directly impede progress in intraoperative perception, surgical workflow understanding, and autonomous robotic decision-making.


Surgical video generation has emerged as a promising solution. By synthesizing realistic surgical scenes, generative models can augment scarce datasets, enable realistic surgical simulation, and provide rich training environments for robotic policy learning. However, as the field has grown rapidly between 2024 and 2026, a clear conceptual framework has been lacking, making it difficult to compare approaches and understand foundational shifts.


2. A Taxonomy of Surgical Video Generation


We categorize the recent literature into three main paradigms:


2.1 Unconditional Generation


Early methods focus on synthesizing visually plausible surgical frames without external input. These models learn the underlying distribution of surgical video data and generate new samples from random noise. While useful for data augmentation, they provide limited control over the content or structure of the generated sequences.


2.2 Conditional Generation


Conditional approaches incorporate explicit inputs—such as anatomical labels, surgical phases, or instrument trajectories—to guide the generation process. This paradigm offers greater controllability, enabling the creation of targeted training samples for specific procedures or scenarios. The trade-off, however, is an increased reliance on annotated data and more complex training pipelines.


2.3 World Model Generation


The most recent development, world model generation, represents a fundamental shift in task definition. Instead of merely synthesizing frames, these models aim to learn the causal dynamics of surgical scenes—predicting how tissue responds to manipulation, how instruments interact with the environment, and how procedures unfold over time. This paradigm moves the field from "looking realistic" to "behaving correctly," a crucial step toward trustworthy surgical simulation and robotic policy learning.


3. Bridging the Gap: Fidelity vs. Clinical Plausibility


Across all categories, a persistent gap remains between pixel-level fidelity and clinical plausibility. While state-of-the-art models can generate visually convincing frames, subtle but clinically critical aspects—such as physiological consistency, tissue mechanical properties, or procedural correctness—often remain inaccurate. Our survey identifies four key bottlenecks:


  • Generalization: Models often fail to generalize beyond the narrow domain of their training data, particularly across different surgical specialties or hospital protocols.
  • Physical realism: Ensuring physically correct tissue deformation, fluid dynamics, and tool–tissue interactions remains a major challenge.
  • Controllability: Achieving fine-grained control over generated content without compromising quality is still an open problem.
  • Interpretability: Understanding what the model has learned and why it makes certain generation choices is crucial for clinical acceptance and regulatory approval.

4. Quantitative Evaluation on Public Datasets


To provide a concrete reference for the community, we summarize experimental results of representative methods on widely used public datasets. This quantitative comparison highlights current strengths and weaknesses, and it establishes a baseline for future work in this rapidly evolving field.


5. Conclusion


The field of surgical video generation is transitioning from a focus on visual realism to a deeper modeling of clinical dynamics. This survey provides a structured overview of this evolution, categorizing the 2024–2026 literature into unconditional generation, conditional generation, and world model generation. By identifying the critical bottlenecks of generalization, physical realism, controllability, and interpretability, we aim to guide future research toward more clinically applicable solutions. For researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science, this survey serves as both a reference point and a call to action to bridge the gap between what is visually plausible and what is clinically plausible.




For citation: Huang, F., Zhang, C., Han, L., & Zhang, L. (2026). Surgical Video Generation: From Diffusion to World Models — A Survey. arXiv:2608.26214 [cs.CV]. https://doi.org/10.48550/arXiv.2608.26214

via ArXiv CV

Related