Establishing the robustness metric as a scalable proxy for clinical relevance in medical AI explainability.
Authors
Affiliations (3)
Affiliations (3)
- Department of Artificial Intelligence, Korea University, Republic of Korea.
- Department of Radiology, Anam Hospital, Korea University College of Medicine, Republic of Korea.
- Department of Artificial Intelligence, Korea University, Republic of Korea; Department of Brain and Cognitive Engineering, Korea University, Republic of Korea. Electronic address: [email protected].
Abstract
Despite high performance of deep learning in medical imaging applications, the critical lack of validated computational metrics for explainable AI (XAI) impedes clinical integration. To address this gap, our study introduces a multi-level validation framework to rigorously assess seven computational evaluation metrics applied to eight widely-used post-hoc attribution methods - spanning gradient-based, input attribution, and decomposition-based families - on large-scale structural MRI datasets (UK Biobank and ADNI) with two different benchmark tasks across several deep learning architectures. At the first level, we benchmark metrics against three scales of clinical relevance: voxel-based morphometry, regional volumetric associations, and expert radiologist assessments, revealing that many widely used conventional metrics exhibit negligible correlation with clinical evidence. At the second level, we show that metrics in addition are often significantly confounded by model architecture rather than reflecting explanation quality. At the third level, we conduct a computational efficiency and stability analysis focused on practical dimensions of the metrics. Across all levels, our statistical analyses show that only one of the metrics has consistent high performance: the Robustness score - defined as the stability of explanations across random training initializations - demonstrates strong alignment with morphometry, volumetric associations, and human expert judgments, effectively isolates the quality of the XAI method from architectural bias, and possesses good computational efficiency. Our novel, multi-level framework therefore establishes Robustness as a superior, scalable proxy for clinical validity and trustworthy AI in medical imaging.