Artificial intelligence for skeletal classification in orthodontics: A systematic review and meta-analysis of lateral cephalometric studies
Journal of Taibah University Medical Sciences, cilt.21, sa.4, ss.683-697, 2026 (ESCI, Scopus)
- Yayın Türü: Makale / Derleme
- Cilt numarası: 21 Sayı: 4
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.jtumed.2026.06.003
- Dergi Adı: Journal of Taibah University Medical Sciences
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, EMBASE, Directory of Open Access Journals
- Sayfa Sayıları: ss.683-697
- Anahtar Kelimeler: Artificial intelligence, Cephalometry, Deep learning, Diagnosis, Orthodontics
- Ankara Üniversitesi Adresli: Evet
Özet
Objective To synthesize the available evidence on the diagnostic performance of artificial intelligence for sagittal skeletal classification using lateral cephalometric radiographs and to identify methodological limitations affecting clinical translation. Methods PubMed, Embase, Scopus, and Web of Science were searched up to July 13, 2025, following PRISMA guidelines. Eligible studies used artificial intelligence to classify skeletal Class I, II, and III relationships from lateral cephalograms and reported diagnostic performance metrics. Non-image-based models, landmarking-only studies, commercial software evaluations, and non-original reports were excluded. Risk of bias was assessed using QUADAS-2. Exploratory random-effects meta-analyses using restricted maximum likelihood were performed, with subgroup analyses by skeletal class. Because of the small number of eligible studies, heterogeneity, and multiple model-level results from shared data sets, pooled estimates were interpreted cautiously. Results Ninety-one records were screened, 12 underwent full-text review, four met the inclusion criteria, and three were included in quantitative synthesis, comprising an effective test sample of 2,743 across nine model arms. One study was excluded from pooling because it used combined posteroanterior and lateral inputs with best-fold reporting, limiting comparability. The pooled multiclass performance showed sensitivity of 87.0%, specificity of 91.1%, accuracy of 85.3%, and area under the receiver operating characteristic curve of 0.94. DenseNet-based models generally showed the strongest performance, whereas the Swin-T transformer model showed the weakest performance. QUADAS-2 indicated low to moderate risk of bias, with major concerns related to single-center sampling and lack of external validation. Conclusions Artificial intelligence shows promising diagnostic performance for sagittal skeletal classification from lateral cephalograms. However, the current evidence remains limited, heterogeneous, and mainly single center. Therefore, pooled estimates should be considered exploratory rather than definitive indicators of clinical performance.