Hybrid Machine Learning Model for Phishing Detection Using Multi-Modal Data
Epelle Bright Curthbert
Source record
Source: Crossref
Published: Sep 17, 2026
DOI: 10.56201/ijcsmt.vol.12.no2.2026.pg258.264
Open original source ↗Source abstract
Phishing attacks continue to pose a severe threat to cybersecurity, exploiting deceptive websites and communications to steal sensitive user information. Traditional detection methods relying on single-modality data, such as URLs or heuristics, often fall short against sophisticated and zero-day attacks. This study proposes a hybrid machine learning model that leverages multi-modal data—including textual URL features, HTML content, webpage screenshots, and metadata—to achieve robust and accurate phishing detection. The model integrates deep learning architectures (e.g., CNNs for visual features, BiLSTM and BERT for textual sequences, and gated MLPs) with ensemble techniques for late-fusion of modalities, enhanced by explainable AI (XAI) methods like SHAP for interpretability. Evaluated on large-scale benchmark datasets, the proposed framework demonstrates superior performance, attaining accuracies up to 99.83%, high F1-scores, and strong resilience to evolving threats. Results highlight the efficacy of multi-modal fusion in overcoming limitations of unimodal approaches, with implications for real-time deployment in web and mobile environments. This research advances phishing detection by offering a scalable, interpretable, and efficient solution, while identifying avenues for future enhancements in multilingual support and adversarial robustness.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.