The role of large language models as screening assistants in the diagnosis of placenta accreta spectrum pathologies.
Authors
Affiliations (2)
Affiliations (2)
- Division of Gynecologic Ultrasound and Prenatal Diagnostics, Department of Gynecology and Obstetrics, University Hospital Basel, Basel, Switzerland.
- Department of Pathology, Nuremberg Clinic, Paracelsus Medical University, Nuremberg, Germany.
Abstract
Large language models (LLMs) are increasingly utilized in modern medicine. In obstetrics, their ability to interpret complex imaging, such as ultrasound images for placenta accreta spectrum (PAS), remains largely unexplored. This study aimed to investigate the capacity of multimodal LLMs to interpret sonographic images of placentae and differentiate between normal findings and various types of PAS. In a comparative, two-run, cross-sectional design, three LLMs (Chat-GPT, Google Gemini, and Claude) were tested using 18 cases of PAS and 18 normal placental ultrasound images (B mode and Doppler). Four clinical experts evaluated the models' performance, diagnostic accuracy, and justification of findings using standardized evaluation criteria. Statistical analyses, including Cochran's Q and McNemar's tests, were used to compare model performance. Gemini and Chat-GPT significantly outperformed Claude in diagnostic accuracy (P < 0.001). Gemini demonstrated the highest sensitivity (100%), followed by Chat-GPT (88.9%), whereas Claude's sensitivity was notably low (5.6%). However, all models exhibited poor specificity, with Gemini reaching only 33.3% and both Chat-GPT and Claude at nearly 0%, frequently misidentifying normal placentae as abnormal. Gemini and Chat-GPT show high sensitivity for detecting PAS but are limited by low specificity and a tendency to report non-standard findings. Currently, LLMs cannot function as independent diagnostic tools in obstetric imaging for PAS but may serve as high-sensitivity screening assistants for triage to specialized centers.