Analysis and Development of Transformer (DistilBERT)-LSTM and Reinforcement Learning Models for Adaptive Phishing Email Detection

Farizal Herry Saputra, Kahfi Heryandi Suradiraja, Abu Khalid Rivai

Abstract


Phishing detection faces two major challenges: performance degradation caused by domain shift and a heavy reliance on costly labeled data. This study proposes an adaptive phishing email detection model that integrates a hybrid DistilBERT–LSTM architecture with a Proximal Policy Optimization (PPO)-based Reinforcement Learning agent. The proposed methodology employs a multi-stage transfer learning framework using three datasets: Enron as the source domain, Phishing_Validation for supervised domain adaptation, and CEAS_08 to simulate an unlabeled data stream through pseudo-labeling. Experimental results demonstrate excellent performance on the source-domain dataset (Enron), achieving an F1-score of 0.9935. However, the model's performance declined on the Phishing_Validation dataset (F1-score = 0.9067), confirming the impact of domain shift. By incorporating the PPO agent, the proposed model autonomously recovered its performance on the CEAS_08 dataset, achieving an F1-score of 0.9516, an accuracy of 0.9468, and a ROC–AUC of 0.9915. The stability of the adaptation process was validated by the convergence of the Kullback–Leibler (KL) divergence to 0.000173, although a minor overconfidence of approximately 5% was observed between the model's confidence estimates and the ground-truth labels. These findings demonstrate the effectiveness of PPO in mitigating domain shift within an unsupervised adaptation environment. Future research should focus on improving model calibration and exploring multimodal feature integration to further strengthen cybersecurity defenses.

Keywords


DistilBERT-LSTM; domain shift; phishing detection; proximal policy optimization (PPO); pseudo-labeling

Full Text:

PDF

References


P. Sari and T. Sutabri, “Analisis Kejahatan Online phising pada Institusi Pemerintah/Pendidik Sehari-hari,” J. Digit. Teknol. Inf., Vol. 6, No. 1, p. 29, 2023, DOI: 10.32502/digital.v6i1.5620.

S. Salloum, T. Gaber, S. Vadera, and K. Shaalan, “A Systematic Literature Review on Phishing Email Detection using Natural Language Processing Techniques,” IEEE Access, Vol. 10, pp. 65703–65727, 2022, DOI: 10.1109/ACCESS.2022.3183083.

P. H. Kyaw, J. Gutierrez, and A. Ghobakhlou, “A Systematic Review of Deep Learning Techniques for Phishing Email Detection,” Artif. Intell. Rev., Vol. 13, No. 3823, 2024, DOI: 10.1007/s10462-024-10944-7.

Id-SIRTII /CC, “Lanskap Keamanan Siber Indonesia 2024,” 2023.

I. A. Fares et al., “Deep Transfer Learning based on Hybrid Swin Transformers with LSTM for Intrusion Detection Systems in IoT Environment,” IEEE Open J. Commun. Soc., Vol. 6, No. April, pp. 4342–4365, 2025, DOI: 10.1109/OJCOMS.2025.3569301.

N. Marshall, D. Sturman, and J. C. Auton, “Exploring the Evidence for Email Phishing Training: A Scoping Review,” Comput. Secur., Vol. 139, No. November 2023, p. 103695, 2024, DOI: 10.1016/j.cose.2023.103695.

M. R. Naeem, R. Amin, M. Farhan, F. S. Alsubaei, E. Alsolami, and M. D. Zakaria, “Cyber Security Enhancements with Reinforcement Learning: A Zero-Day Vulnerabilityu Identification Perspective,” PLoS One, Vol. 20, No. 5 May, pp. 1–33, 2025, DOI: 10.1371/journal.pone.0324595.

N. Tatipatri and S. L. Arun, “A Comprehensive Review on Cyber-Attacks in Power Systems: Impact Analysis, Detection, and Cyber Security,” IEEE Access, Vol. 12, No. February, pp. 18147–18167, 2024, DOI: 10.1109/ACCESS.2024.3361039.

B. Kommey, O. J. Isaac, E. Tamakloe, and D. Opoku4, “Reinforcement Learning Review: Past Acts, Present Facts and Future Prospects,” IT J. Res. Dev., Vol. 8, No. 2, pp. 120–142, 2024, DOI: 10.25299/itjrd.2023.13474.

K. Thakur, M. L. Ali, M. A. Obaidat, and A. Kamruzzaman, “A Systematic Review on Deep-Learning-based Phishing Email Detection,” Electron., Vol. 12, No. 21, pp. 1–26, 2023, DOI: 10.3390/electronics12214545.

Y. R. Purnamadewi and A. Zahra, “Enhancing Detection of Zero-Day Phishing Email Attacks in the Indonesian Language using Deep Learning Algorithms,” Bull. Electr. Eng. Informatics, Vol. 14, No. 1, pp. 505–512, 2025, DOI: 10.11591/eei.v14i1.8759.

F. Akhbardeh, M. Zampieri, C. O. Alm, and T. Desell, “Transfer Learning Methods for Domain Adaptation in Technical Logbook Datasets,” 2022 Lang. Resour. Eval. Conf. Lr. 2022, No. June, pp. 4235–4244, 2022.

F. A. Shaikh and M. Siponen, “Organizational Learning from Cybersecurity Performance: Effects on Cybersecurity Investment Decisions,” Inf. Syst. Front., 2023, DOI: 10.1007/s10796-023-10404-7.

Z. Alshingiti, R. Alaqel, J. Al-Muhtadi, Q. E. U. Haq, K. Saleem, and M. H. Faheem, “A Deep Learning-based Phishing Detection System using CNN, LSTM, and LSTM-CNN,” Electron., Vol. 12, No. 1, pp. 1–18, 2023, DOI: 10.3390/electronics12010232.

S. Dey, W. Sarma, and S. Tiwari, “AI-Powered Phishing Detection : Integrating Natural Language Processing and Deep Learning for Email Security,” Vol. 10, No. 02, pp. 394–415, 2023.

R. Meléndez, M. Ptaszynski, and F. Masui, “Comparative Investigation of Traditional Machine-Learning Models and Transformer Models for Phishing Email Detection,” Electron., Vol. 13, No. 24, 2024, DOI: 10.3390/electronics13244877.

S. Atawneh and H. Aljehani, “Phishing Email Detection Model using Deep Learning,” Electron., Vol. 12, No. 20, 2023, DOI: 10.3390/electronics12204261.

E. Laparra, A. Mascio, S. Velupillai, and T. Miller, “A Review of Recent Work in Transfer Learning and Domain Adaptation for Natural Language Processing of Electronic Health Records,” Yearb. Med. Inform., Vol. 30, No. 1, pp. 239–244, 2021, DOI: 10.1055/s-0041-1726522.

B. Dhyani, “Transfer Learning in Natural Language Processing: A Survey,” Math. Stat. Eng. Appl., Vol. 70, No. 1, pp. 303–311, 2021, DOI: 10.17762/msea.v70i1.2312.

C. Subhadeep, “Phishing Email Dataset,” 2023, kaggle.com. DOI: https://doi.org/10.34740/kaggle/dsv/6090437.

G. Miltchev, R., Dimitar, R., & Evgeni, “Phishing Validation Emails Dataset,” 2024, zenodo. DOI: https://doi.org/10.5281/zenodo.13474746.

H. Syahrizal and M. S. Jailani, “Jenis-Jenis+Penelitian+Dalam+Penelitian+Kuantitatif+dan+Kualitatif,” J. Pendidikan, Sos. dan Hum., Vol. 1, pp. 18–22, 2023, [Online]. Available: https://ejournal.yayasanpendidikandzurriyatulquran.id/index.php/qosim/article/view/49

M. Songailaitė, E. Kankevičiūtė, B. Zhyhun, and J. Mandravickaitė, “BERT-based Models for Phishing Detection,” CEUR Workshop Proc., Vol. 3575, pp. 34–44, 2023.

P. M. Gholampour and R. M. Verma, “Adversarial Robustness of Phishing Email Detection Models,” IWSPA 2023 - Proc. 9th ACM Int. Work. Secur. Priv. Anal., 2023, DOI: 10.1145/3579987.3586567.

A. Bahaa, A. Kamal, H. Fahmy, and A. S. Ghoneim, “DB-CBIL: A DistilBert-based Transformer Hybrid Model using CNN and BiLSTM for Software Vulnerability Detection,” IEEE Access, Vol. 12, No. May, pp. 64446–64460, 2024, DOI: 10.1109/ACCESS.2024.3396410.

M. R. Amini, V. Feofanov, L. Pauletto, L. Hadjadj, É. Devijver, and Y. Maximov, “Self-Training: A Survey,” Neurocomputing, Vol. 616, No. November 2024, p. 128904, 2025, DOI: 10.1016/j.neucom.2024.128904.




DOI: https://doi.org/10.32520/stmsi.v15i7.6432

Article Metrics

Abstract view : 11 times
PDF - 1 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.