Comparison of the diagnostic accuracy of dentists and ChatGPT in jawbone lesions


Satan N. S., PAMUKÇU U., TÜRK B. E., AKARSLAN Z., AÇIK KEMALOĞLU S., PEKER İ.

BMC Oral Health, cilt.26, sa.1, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 26 Sayı: 1
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1186/s12903-026-08623-w
  • Dergi Adı: BMC Oral Health
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Directory of Open Access Journals, Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: Artificial intelligence in healthcare, ChatGPT, Diagnosis, Jawbone Lesion, Large Multimodal Model (LMM), Maxillofacial radiology
  • Ankara Üniversitesi Adresli: Evet

Özet

Background/objectives: Artificial intelligence (AI) is leading to a significant paradigm shift in medical imaging and diagnostic sciences. In particular, Chat Generative Pre-trained Transformer (ChatGPT) is finding increasing application in diagnostic processes due to its ability to generate clinical outcomes. This study aims to evaluate the diagnostic accuracy of ChatGPT for jawbone lesions and also to compare it with that of Oral and Maxillofacial Radiologists (OMFR), Oral and Maxillofacial Surgeons (OMFS), and general dentists. Materials & methods: Thirty cases with jawbone lesions, for which clinical information, panoramic radiographs, and histopathological diagnoses were available, were selected. A questionnaire was prepared, including participants’ (OMFR, OMFS, and general dentists) demographic information, the cases’ clinical findings and panoramic radiographs, and distributed via electronic communication channels. The same cases were loaded into ChatGPT-4 and asked to generate a preliminary diagnosis. The data were statistically analyzed using the Wilcoxon Signed Rank, Mann–Whitney U, and Kruskal–Wallis tests at a significance level of p < 0.05. Results: Overall, ChatGPT’s diagnostic accuracy was limited to 46.67%, while the OMFR (67.71%) and OMFS (58.96%) groups had statistically significantly higher success rates (p < 0.05) than ChatGPT. General dentists (37.85%) had lower or similar diagnostic accuracy compared to ChatGPT in most subgroups (gender, age, workplace, professional experience). Conclusions: ChatGPT demonstrated moderate diagnostic accuracy. While OMFR and OMFS participants had significantly higher accuracy rates than ChatGPT, ChatGPT generally outperformed general dentists. These results indicate that such AI systems cannot replace specialist clinicians but can provide valuable contributions as supportive tools that enhance diagnosis.