Back to all papers

Vision Language Models for Ultrasound Assessment of Suspicious Axillary Lymph Nodes in Breast Cancer.

July 22, 2026pubmed logopapers

Authors

He P,Chen C,Dong HD,Chen MY,Lin ZH,Wu SL,Cai WF,Lin XX,Dong YQ,Wang C,Fu FM

Affiliations (6)

  • Department of Breast Surgery, Fujian Medical University Union Hospital, Fuzhou, Fujian Province, 350001, China.
  • Department of General Surgery, Fujian Medical University Union Hospital, Fuzhou, Fujian Province, 350001, China.
  • Breast Cancer Institute, Fujian Medical University, Fuzhou, Fujian Province, China.
  • Fujian Key Laboratory of Medical Bioinformatics, Fujian Medical University, Fuzhou 350122, China.
  • Department of Ultrasound, Fujian Medical University Union Hospital, Fuzhou 350001, China.
  • Department of Cardiology, Fujian Medical University Union Hospital, Fuzhou, China.

Abstract

<b>Background:</b> Axillary management in patients with breast cancer is becoming increasingly individualized, requiring accurate nodal characterization. <b>Objective:</b> To evaluate the performance of vision-language models (VLMs) with image-only diagnostic input for detecting malignancy in biopsied axillary lymph nodes in patients with breast cancer and to compare performance with radiologists of different experience levels. <b>Methods:</b> This retrospective study included 718 patients (mean age, 53.1±11.4 years) with primary invasive breast cancer who underwent ultrasound-guided fine-needle aspiration or core-needle biopsy (FNA/CNB) of a suspicious target axillary lymph node from January 2025 through June 2025. A grayscale image and corresponding Doppler image, if available, were saved for a single biopsied node in each patient. Two inexperienced radiologists (both with only residency exposure to breast ultrasound) and two experienced radiologists (both breast imaging radiologists) reviewed images to classify nodes as metastatic or nonmetastatic; each pair reached consensus. The images were also inputted to three VLMs (GPT-5.2, Gemini-3-Pro, and Claude-Opus-4.5), with a text prompt requesting binary outputs (metastatic or nonmetastatic). Performance was compared between the reader groups and between the models and reader groups using Bonferroni adjustment and FNA/CNB as reference. <b>Results:</b> Accuracy, sensitivity, and specificity for inexperienced radiologists were 76%, 92%, and 64%; for experienced radiologists were 83%, 75%, and 89%; for GPT-5.2 were 78%, 76%, and 80%; for Gemini-3-Pro were 73%, 79%, and 68%; and for Claude-Opus-4.5 were 67%, 63%, and 70%, respectively. Accuracy was significantly higher for experienced radiologists compared with inexperienced radiologists and all three models; and for inexperienced radiologists compared with Claude-Opus-4.5. Sensitivity was significantly higher for inexperienced radiologists compared with experienced radiologists and all three models and for experienced radiologists compared with Claude-Opus-4.5. Specificity was significantly higher for experienced radiologists compared with inexperienced radiologists and all three models and for GPT-5.2 compared with inexperienced radiologists. Other comparisons were not significant. <b>Conclusion:</b> GPT-5.2 (the VLM with highest accuracy) showed no significant difference in accuracy versus inexperienced radiologists, albeit had lower sensitivity and greater specificity. GPT-5.2 had lower accuracy than experienced radiologists. <b>Clinical Impact:</b> The results support investigation of VLMs as a supervised decision-support tool to help reduce false-positive assessments by inexperienced radiologists.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.