A feature-enriched deep learning based ensemble framework for robust phishing URL detection
PeerJ Computer Science, cilt.12, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 12
- Basım Tarihi: 2026
- Doi Numarası: 10.7717/peerj-cs.4029
- Dergi Adı: PeerJ Computer Science
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, Directory of Open Access Journals, Technology Collection (ProQuest)
- Anahtar Kelimeler: Deep learning, Ensemble, Future engineering, Machine learning, Phishing URL
- Ankara Üniversitesi Adresli: Evet
Özet
Through the prevalence of internet technologies and the integration of digitalization into our daily routine, a significant portion of our personal, financial, and professional activities has shifted to cyberspace. However, this transformation has led to the rise of cyber threats, with phishing attacks being among the most dangerous and widespread, often tricking users into disclosing sensitive information by mimicking legitimate websites. In this study, we introduce a feature-driven framework for phishing Uniform Resource Locator (URL) detection, emphasizing the design and evaluation of enhanced feature representations. The proposed approach integrates structural, lexical, and distributional characteristics of URLs and evaluates their effectiveness across both classical machine learning and deep learning models. To assess their impact, several machine learning algorithms, including Naive Bayes, k-Nearest Neighbors, Random Forest, Gradient Boosting, and Multi-Layer Perceptron, are employed. To ensure a comprehensive evaluation, experiments are conducted on three datasets, including two recent large-scale datasets and a widely used benchmark dataset, enabling the assessment of model performance under different data distributions. Among all machine learning models, Random Forest achieves highly competitive performance with an accuracy of up to 99.82, demonstrating the effectiveness of the proposed feature representation even with computationally efficient models. Building on this foundation, multiple deep learning architectures combining convolutional neural network (CNN)-based sequence modeling with handcrafted features are developed and further enhanced using ensemble strategies. The resulting ensemble model achieves the highest overall performance, with accuracy reaching 99.84 and consistently low false negative rates across all datasets. However, the observed improvements over strong classical baselines remain modest, highlighting that performance gains are primarily driven by feature design rather than model complexity. Besides, we conduct multiple statistical significance tests, including paired t-tests, Wilcoxon signed-rank tests, Analysis of Variance (ANOVA), and McNemar's test, to assess the reliability of the observed performance differences. The results confirm that while the improvements of the proposed approach are statistically consistent, the performance gaps among top-performing models are relatively small in practical terms.