Skip to content
Open access

Adaptive Phishing URL Detection Using Hybrid Fuzzy C-Means Clustering and XGBOOST

Aug 2026 · Karbala International Journal of Modern Science · Vol 12 · 0 citations

TL;DR

A hybrid phishing URL detection system that integrates Fuzzy C-Means clustering with XGBoost classification, enhanced by a novel Micro Adaptive Feature Extractor (MAFE) is proposed.

Abstract

Phishing attacks continue to evolve in sophistication, rendering static detection methods increasingly ineffective. Existing URL-based approaches suffer from limited adaptability to emerging phishing patterns, mislabeled training data, and insufficient validation protocols. This paper proposes a hybrid phishing URL detection system that integrates Fuzzy C-Means (FCM) clustering with XGBoost classification, enhanced by a novel Micro Adaptive Feature Extractor (MAFE). The system employs a multi-stage pipeline: feature engineering generating 36 statistical and interaction features, MAFE producing 15 adaptive features through class-aware dynamic weighting, micro-pattern detection, and entropy analysis, and FCM with K=2 clusters providing soft membership features to XGBoost. A two-pass confidence-based mislabel detection protocol identifies and removes 2.66% suspected labeling errors from the training data. The system is evaluated on the large-scale DEPHIDES dataset of 5,202,841 URLs using a proper three-way split: 60% training, 10% validation, and 30% test. The classification threshold is optimized exclusively on the validation set, ensuring unbiased test evaluation. The proposed system achieves 97.86% accuracy and 99.84% AUC on the raw test set, improving to 98.98% accuracy after verified mislabel removal. Comparative evaluation demonstrates that the system outperforms Random Forest 95.91%, LightGBM 96.31%, CatBoost 95.75%, and standalone XGBoost 96.59% trained on identical data with the same evaluation protocol. The system processes URLs at 6,528 URLs/second, with 95% confidence intervals of 97.77%–97.96% for accuracy. A sensitivity analysis confirms robustness to the MAFE adaptation rate parameter, with accuracy varying by only 0.06% across α ∈ [0.05, 0.30].

Read PDF

Similar papers

Open access Aug 2026

Real-Time Phishing URL Detection Using a Hybrid Stacking Ensemble: Gradio and Browser Extension Deployment

Phishing attacks remain a prevalent and rapidly evolving cybersecurity threat, leveraging deceptive Uniform Resource Locators (URLs) and fraudulent websites to steal sensitive user data, financial credentials, and personal information. Traditional detection mechanisms, such as blacklist-based and heuristic approaches,...

E. Kavya, A. S. Chakravarthy · 0 citations
Aug 2026

Distributed phishing URL classification: Leveraging modified XGBoost in network environments

A distributed phishing URL classification algorithm, called modified XGBoost (MD-XGBoost), is presented to bridge the gap existing between deep learning (DL) based high-accuracy but computationally-intensive methods and interpretable and computationally-efficient machine learning models to use in practice on a distribu...

G. R, G. S., Belshia Jebamalar T et al. · 0 citations
Open access Sep 2026

URL-based phishing detection using XGBoost with engineered features

URL-based phishing involves fake uniform resource locators (URLs) created by attackers to trick users into believing they are visiting a legitimate website and thereby steal their confidential information. While several powerful machine learning (ML) and deep learning (DL) studies exist to detect phishing, they still f...

Jawaher Alharbi, M. Bayousef, Hind Almisbahi · 0 citations
Open access Sep 2026

Real-Time Phishing Attack Detection in A Chrome Extension Using Hybrid Gwo-Pso Feature Selection and Optimized Xgboost

Findings demonstrate that hybrid metaheuristic optimization combined with gradient-boosted classification provides an efficient, lightweight, and empirically verifiable solution for real-time phishing defense at the browser edge.

A. Giwa-Raheem, Dahiru Haliru Ibrahim, Abubakar Faruku Saad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.