Back to all papers

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

July 24, 2026pubmed logopapers

Authors

Akdogan AI,Akdogan EK,Tumer MF,Kutlu S,Tekindal MA,Tosun O

Affiliations (4)

  • Department of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey. [email protected].
  • Department of Orthopedics and Traumatology, Bakircay University Cigli Training and Research Hospital, Izmir, Turkey.
  • Department of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey.
  • Department of Biostatistics, Izmir Katip Celebi University, Izmir, Turkey.

Abstract

To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren-Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability. In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0-1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures. Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1-2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485). Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.