Can we really trust generative models in healthcare? A systematic review of uncertainty quantification in generative AI for medical imaging.
Authors
Affiliations (7)
Affiliations (7)
- Department of Electronics and Telecommunications, Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129 Turin, Italy.
- National Metrology Institute of Italy (INRiM), Turin, Italy.
- Department of Electronics and Telecommunications, Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129 Turin, Italy; National Metrology Institute of Italy (INRiM), Turin, Italy.
- Manipal Institute of Technology, Manipal Academy of Higher Education, Manipal, India.
- Electrical-Electronics Engineering Department, Technology Faculty, Firat University, Elazig, Turkey.
- Centre for Health Research, University of Southern Queensland, Australia; School of Mathematics, Physics and Computing, University of Southern Queensland, Springfield, Australia.
- Department of Electronics and Telecommunications, Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129 Turin, Italy. Electronic address: [email protected].
Abstract
Generative AI is transforming medical imaging research through synthesis, enhancement, and reconstruction of clinical images. While these advances show promise in addressing data scarcity and supporting diagnostics, clinical adoption remains limited due to challenges in assessing the trustworthiness of generated content. This study aims to systematically evaluate the integration of uncertainty quantification (UQ) methods within generative models for medical imaging to enhance result reliability. A systematic review following PRISMA guidelines was conducted, analyzing studies from January 2018 to December 2025 that combined generative models with UQ techniques in medical imaging. The search strategy covered major medical and computer science databases, with studies evaluated against predefined inclusion criteria focusing on implementation methodology and performance metrics. From the analysis of 41 eligible studies, 29 focused on radiology, 8 on microscopy, and 4 on optical coherence tomography. Across this heterogeneous body of evidence, integrating UQ was frequently associated with improved performance or with more informative reliability assessment, including reported gains in reconstruction quality, segmentation accuracy, and anomaly detection. Notably, 56% of studies (n=23) were published in 2025, indicating rapid field growth. UQ integration represents a crucial advancement toward trustworthy generative AI systems in medical imaging. Key priorities identified include standardizing uncertainty metrics, developing computationally efficient frameworks, and embedding uncertainty awareness within generation processes. These findings suggest that UQ methods can enhance the clinical reliability of generative AI applications in medical imaging.