Back to all papers

Automated Quality Control of Breast Ultrasound Reports Using a BI-RADS-Prompted Large Language Model: A Pilot Multicenter Study Involving 60 Hospitals in China.

July 22, 2026pubmed logopapers

Authors

Lu S,Wang H,Wang J,Ye Y,Guo X,Ni D,Jiang Y

Affiliations (5)

  • Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China; National Ultrasound Medical Quality Control Center, Beijing, China.
  • Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China; National Ultrasound Medical Quality Control Center, Beijing, China; Deputy Director and Secretary General, National Ultrasound Quality Control Center, and Office Director, Beijing Ultrasound Quality Control Center. Electronic address: [email protected].
  • College of Computer Science and Software Engineering, Shenzhen University,Shenzhen, China; Medical Ultrasound Image Computing Lab, School of Artificial Intelligence, Shenzhen University, Shenzhen, China.
  • Medical Ultrasound Image Computing Lab, School of Artificial Intelligence, Shenzhen University, Shenzhen, China; National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University, Shenzhen, China.
  • Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China; National Ultrasound Medical Quality Control Center, Beijing, China; Director, National Ultrasound Medical Quality Control Center.

Abstract

To evaluate the feasibility of using large language model (LLM) for automated quality control (QC) of breast ultrasound (US) reports. In this retrospective multicenter study, we collected breast US reports from 735 patients who had pathological confirmation of mass-type lesions on breast US across 60 hospitals in China. Each hospital's QC personnel converted free-text reports into standardized structured outputs. A gold standard was established through a multilevel expert review. The Qwen2.5-VL-7B LLM was applied to the same free-text reports, generating structured reports based on BI-RADS prompts. Accuracy was compared both between the LLM and QC personnel as well as the times required to produce the outputs for analysis of efficiency. The LLM demonstrated higher accuracy in conducting QC for key breast lesion US features, such as margin (80.0% versus 64.1%, P < .0001) and echo pattern (74.0% versus 56.1%, P < .0001). Subgroup analysis further confirmed its robustness: in the complex reports of multiple lesions, it maintained its advantage in margin QC (79.8% versus 64.4%). The study also revealed a significant positive correlation between the LLM's QC accuracy and BI-RADS categories (from 3 to 5) (Spearman ρ = 0.264, P = .028), a trend not observed in manual QC. In terms of efficiency, the LLM completed QC for 50 reports in an average of only 13 min, faster than 212.5 min required by manual reviewers. The proposed LLM-based system provides a reliable, accurate, and efficient solution for breast US reports. It achieves human-comparable performance with markedly higher efficiency, particularly in reports with high suspicion levels, offering a reliable tool to enhance report quality.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.