Task-Specific NLP Outperforms General LLMs for Lung Nodule Detection in Chest CT Reports
A radiology-trained NLP model significantly outperformed several general-purpose large language models in extracting incidental lung nodule data from chest CT reports.
Key Details
- 1Task-specific NLP model ('FiNd') was compared with seven general-purpose LLMs (Gemma, Haiku, Sonnet 2, GPT-4o, DeepSeek, Phi-4, MedGemma).
- 2FiNd was developed using 21,542 radiology reports and tested on 1,016 chest CT reports.
- 3Performance was assessed for identifying incidental lung nodules (ILNs) and categorizing them by size (<6 mm, 6–7.9 mm, ≥8 mm).
- 4FiNd achieved 96.8% accuracy for nodules ≥6 mm and 97.4% for nodules ≥8 mm, outperforming all general LLMs.
- 5General-purpose LLM accuracy ranged from 77.7% to 88.6% for detection tasks; specificity varied widely among models.
- 6Researchers urge adaptation of LLMs using radiology-specific training for better clinical integration.
Why It Matters

Source
AuntMinnie
Related News

AI Models Reveal Racial Disparities in Breast Cancer Patterns
Machine learning models reveal significant racial disparities and key predictors in breast cancer incidence across diverse groups.

AI Algorithm Streamlines and Standardizes Shoulder Ultrasound Acquisition
A multitask AI system demonstrated high accuracy in standardizing and guiding shoulder musculoskeletal ultrasound imaging.

Deepfake X-rays Fool Radiologists and AI, Raising Security Concerns
Both radiologists and AI models struggle to differentiate between authentic and AI-generated ('deepfake') radiographic images, raising major security and clinical concerns.