Back to all papers

Closing the Performance Gap between Generalists and Breast Imaging Specialists Using a Nationally Deployed AI Workflow for Screening Mammography.

July 21, 2026pubmed logopapers

Authors

McCabe MP,Wakelin EA,Louis LD,Ng AY,Buist DSM,Lee CI,Sorensen AG,Haslam B

Affiliations (3)

  • DeepHealth, 212 Elm St, Somerville, MA 02144.
  • Data-driven Strategies for Medicine & Biotechnology, Mercer Island, Wash.
  • Department of Radiology, University of Wisconsin-Madison School of Medicine and Public Health, Madison, Wis.

Abstract

Background Radiologist expertise plays a critical role in breast cancer screening outcomes; yet approximately 70% of screening mammography interpretations in the United States are performed by general radiologists (generalists) rather than breast imaging specialists (specialists). Purpose To evaluate the impact of a multistage artificial intelligence (AI)-driven workflow on the clinical performance of generalists and fellowship-trained specialists. Materials and Methods This prospective study included screening mammogram interpretations from radiologists across 109 U.S. imaging facilities performed between September 2021 and December 2022. Only radiologists who interpreted digital breast tomosynthesis screening examinations during both study periods were included, and only bilateral examinations from women aged 35 years or older were included. The multistage AI-driven workflow integrated a computer-aided detection and diagnosis device and a separate "Safeguard Review" that routes AI-identified suspicious screening mammogram examinations that were not recalled for additional expert review. Adjusted cancer detection rate (CDR), positive predictive value (PPV) of recalls, and recall rate (RR) were compared before and after introduction of the multistage AI-driven workflow using logistic regression with generalized estimating equations. Results A total of 95 radiologists (60 generalists with a median of 19 years of experience [IQR, 14-31 years; range, 3-50 years] and 35 specialists with a median of 11 years of experience [IQR, 6-21 years; range, 1-32 years]) interpreting 577 742 examinations were included. Adjusted results show that the CDR of generalists increased from 3.76 (95% CI: 3.46, 4.08) to 4.99 (95% CI: 4.51, 5.53; <i>P</i> < .001) cancers per 1000 examinations with the AI workflow. The CDR of specialists was similar before and after adding the AI workflow (from 4.47 to 4.76 per 1000 examinations, respectively; <i>P</i> = .33) and was similar to that of generalists with AI (<i>P</i> = .53). Among generalists, PPV of recalls increased by 15.09% (from 3.38% [95% CI: 3.03, 3.75] to 3.89% [95% CI: 3.39, 4.46]; <i>P</i> = .02), indicating more efficient cancer detection (more cancers detected per recall) despite a 14.79% relative increase in RR (from 9.06% to 10.40%, <i>P</i> = .007). Specialists showed no change in PPV of recalls with the AI workflow (<i>P</i> = .86). Conclusion A multistage AI-driven workflow was associated with substantially improved CDR and PPV of recalls for generalists, which were on par with those of specialists. © RSNA, 2026 <i>Supplemental material is available for this article.</i> See also the editorial by Schiaffino and Cozzi in this issue.

Topics

MammographyWorkflowBreast NeoplasmsArtificial IntelligenceRadiologistsClinical CompetenceEarly Detection of CancerJournal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.