Back to all papers

Moderate-to-substantial agreement of ChatGPT-5 for Kellgren-Lawrence grading on synthetic knee radiographs: a controlled cross-sectional observer agreement study.

July 20, 2026pubmed logopapers

Authors

Kılınç Ö,Altunel Kılınç E,Çabuk Çelik N

Affiliations (3)

  • Department of Orthopedics and Traumatology, Mersin City Training and Research Hospital, Mersin, Turkey.
  • Department of Rheumatology, Mersin City Training and Research Hospital, Mersin, Turkey. [email protected].
  • Department of Rheumatology, Sincan Training and Research Hospital, Ankara, Turkey.

Abstract

AI models are increasingly explored for radiographic assessment of knee osteoarthritis, but their reliability as an experimental model for KL grading remains uncertain. This study evaluated ChatGPT-5 agreement with expert consensus for KL grading and compared its performance with a secondary AI model. This controlled cross-sectional observer agreement analysis used 630 hand-compiled synthetic posteroanterior fixed flexion knee radiographs selected for technical adequacy, anatomical suitability, and interpretability. Three experienced clinicians, one orthopedic surgeon and two rheumatologists, independently graded all images, and a three-rater majority-consensus expert reference standard was established. ChatGPT-5 assessed the same images using a standardized prompt and image-level protocol. Gemini 2.5 Flash assessed a 112-image subset. Agreement was analyzed using weighted Cohen's κ, Gwet's AC2, and mean absolute error (MAE). Binary classification behavior was evaluated at the KL ≥ 2 threshold. Inter-rater agreement before consensus formation was high: Fleiss' κ was 0.884, Gwet's AC2 was 0.921, and pairwise weighted κ ranged from 0.902 to 0.929. Complete three-rater agreement was observed in 520 radiographs (82.5%), and only 14 cases (2.2%) required adjudication. Compared with the expert reference standard, ChatGPT-5 showed moderate-to-substantial agreement, with weighted κ = 0.680 (95% CI 0.616-0.737), MAE = 0.589, and 60.9% exact five-grade agreement. Exact agreement was highest for KL 0 (81.5%) and KL 4 (66.9%) and lower for intermediate grades. For binary KL ≥ 2 classification, ChatGPT-5 correctly classified 513/630 radiographs, with 81.4% accuracy (95% CI 78.2-84.3), 82.1% sensitivity, 80.3% specificity, 87.6% PPV, and 72.6% NPV. A small but significant shift toward higher grading was observed (p = 0.013), with up-grading more frequent in KL 0-1 and down-grading more frequent in KL 3-4. In the 112-image exploratory subset, Gemini 2.5 Flash showed limited agreement with the expert reference standard (weighted κ = 0.23, MAE = 0.66, exact agreement = 50.0%) and poor agreement with ChatGPT-5 (weighted κ = 0.06). For KL ≥ 2 classification, Gemini achieved 87.5% accuracy, 100% sensitivity, 0% specificity, and 87.5% PPV; NPV was not estimable because no KL < 2 classifications were generated. In this controlled synthetic radiograph study, ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions. However, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.

Topics

Osteoarthritis, KneeKnee JointJournal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.