OCR-mediated modality dominance in vision-language models: implications for radiology AI trustworthiness.
Authors
Affiliations (11)
Affiliations (11)
- Hacettepe University Faculty of Medicine, Department of Pediatric Intensive Care Medicine, Ankara, Turkey.
- Hacettepe University, The Center for Life Support Practice and Research, Ankara, Turkey.
- Ankara Yildirim Beyazit University, Faculty of Medicine, Ankara, Turkey.
- Hacettepe University Faculty of Medicine, Department of Pediatric Emergency Medicine, Ankara, Turkey.
- Middle East Technical University, Informatics Institute, Department of Cognitive Science, Ankara, Turkey.
- Atılım University, School of Medicine, Department of Emergency Medicine, Ankara, Turkey.
- Beth Israel Deaconess Medical Center, Radiology Department, Boston, Massachusetts, United States.
- Massachusetts Institute of Technology, Laboratory for Computational Physiology, Cambridge, Massachusetts, United States.
- Beth Israel Deaconess Medical Center, , Division of Pulmonary, Critical Care and Sleep Medicine, Boston, Massachusetts, United States.
- Harvard T.H. Chan School of Public Health, Department of Biostatistics, Boston, Massachusetts, United States.
- Case Western Reserve University, Department of Computer and Data Sciences, Cleveland, Ohio, United States.
Abstract
Vision-language models (VLMs) are increasingly proposed for radiologic decision support, yet the security implications of deploying optical character recognition (OCR)-capable models in diagnostic workflows remain poorly characterized. When image-embedded text is not treated as untrusted input, the visual channel becomes vulnerable to adversarial manipulation. Ten VLMs, none validated for clinical diagnosis, were evaluated on 600 brain magnetic resonance imaging studies for binary tumor detection under 5 conditions: clean input, visible radiology report injection, human-imperceptible stealth OCR injection, and multi-stage immune-prompt defense. A total of 27,000 inference calls were analyzed. At baseline, performance was heterogeneous, with a median accuracy of 0.69, a sensitivity of 0.79, and a specificity of 0.59. Visible injection caused universal specificity collapse to 0.00 across all models with a false-positive rate (FPR) of 1.00 and a median attack success rate (ASR) of 0.97. Stealth injection, despite being imperceptible to human reviewers, drove substantial degradation with a median accuracy of 0.43, an ASR of 0.57, and an FPR of 0.84. Immune prompting achieved only partial mitigation: under stealth injection, median ASR decreased to 0.44, and accuracy improved to 0.56, yet residual overcalling persisted with a median FPR of 0.67, and three models maintained an FPR of 1.00. Commercial VLMs exhibit a deployment-critical failure mode in radiology-like scenarios: OCR-readable text embedded in images can override pixel-level evidence, even under stealth conditions that evade human inspection. Prompt-level defenses provide insufficient protection. Any clinical integration of VLMs must be governed by system-level safeguards, including OCR-aware input handling, provenance controls, and enforced human verification, before deployment in safety-sensitive environments.