Spam Email Classification Using TF-IDF and Classical Machine Learning on the Enron-Spam Corpus
DOI:
https://doi.org/10.31185/wjcms.562Keywords:
Spam Detection, Malicious Email Classification, Machine Learning, Text Classification, Naive Bayes, Support Vector Machine, Logistic RegressionAbstract
Unsolicited email is a persistent operational and security burden and many high-performing neural approaches require computational and deployment costs that are not necessary for resource-constrained filtering systems. This paper tries to fill this gap by providing a rigorous and reproducible comparison of lightweight classifiers to discriminate spam from legitimate email without overestimating the corpus as a dedicated phishing or malware benchmark. Using 30,089 cleaned and de-duplicated messages from the Enron-Spam corpus, word-level unigram and bigram term frequency–inverse document frequency (TF-IDF) features were evaluated with Multinomial Naive Bayes, Logistic Regression, Linear Support Vector Machine, and Random Forest classifiers. Hyperparameters were selected by grid search on the training partition; stability was assessed through repeated stratified five-fold cross-validation across five random seeds; and paired Logistic Regression and Linear SVM predictions were compared using an exact two-sided McNemar test. On the held-out test set, Logistic Regression achieved the highest accuracy (99.15%) and F1-score (99.13%), followed by Linear SVM (99.05% accuracy; 99.02% F1). Naive Bayes achieved 98.59% accuracy, whereas Random Forest reached 96.69% despite a spam recall of 99.55%. Repeated cross-validation produced identical rounded mean accuracy and macro-F1 for Logistic Regression and Linear SVM (99.08% ± 0.11%), and their held-out difference was not statistically significant (p = 0.2101). These findings support tuned linear TF-IDF models as accurate, stable, and computationally practical spam-filtering baselines, while cross-corpus validation remains necessary.
Downloads
References
[1] Klimt, B., & Yang, Y. (2004). The Enron corpus: A new dataset for email classification research. In Machine Learning: ECML 2004 (LNCS vol. 3201, pp. 217-226). Springer. https://doi.org/10.1007/978-3-540-30115-8_22
[2] Metsis, V., Androutsopoulos, I., & Paliouras, G. (2006). Spam filtering with Naive Bayes - which Naive Bayes? In Proceedings of the 3rd Conference on Email and Anti-Spam (CEAS), Mountain View, CA, USA.
[3] Sahami, M., Dumais, S., Heckerman, D., & Horvitz, E. (1998). A Bayesian approach to filtering junk e-mail. In Proceedings of the AAAI-98 Workshop on Learning for Text Categorization, Madison, WI, USA, pp. 55-62.
[4] Androutsopoulos, I., Koutsias, J., Chandrinos, K. V., Paliouras, G., & Spyropoulos, C. D. (2000). An evaluation of Naive Bayesian anti-spam filtering. In Proceedings of the Workshop on Machine Learning in the New Information Age, ECML 2000, Barcelona, Spain, pp. 9-17.
[5] Yusupov, K., Islam, M. R., Muminov, I., Sahlabadi, M., & Yim, K. (2025). Comparative analysis of machine learning and deep learning models for email spam classification using TF-IDF and word embedding techniques. In Advances on Broad-Band Wireless Computing, Communication and Applications (BWCCA 2024), LNDECT vol. 231. Springer. https://doi.org/10.1007/978-3-031-76452-3_11
[6] Alhuzali, A., Alloqmani, A., Aljabri, M., & Alharbi, F. (2025). In-depth analysis of phishing email detection: Evaluating the performance of machine learning and deep learning models across multiple datasets. Applied Sciences, 15(6), 3396. https://doi.org/10.3390/app15063396
[7] A machine learning and NLP-based approach for efficient email spam detection using TF-IDF and logistic regression. (2026). International Journal of Computer Techniques, 13(3), IJCT-V13I3P22. https://ijctjournal.org/efficient-email-spam-detection-tf-idf/
[8] Al Tawil, A., Almazaydeh, L., Qawasmeh, D., Qawasmeh, B., Alshinwan, M., & Elleithy, K. (2024). Comparative analysis of machine learning algorithms for email phishing detection using TF-IDF, Word2Vec, and BERT. Computers, Materials & Continua, 81(2), 3395-3412. https://doi.org/10.32604/cmc.2024.057279
[9] Alrammahi, A. A. H., et al. (2025). Enhancing spam detection with advanced feature extraction and unsupervised clustering. International Journal of Information Technology. Springer Nature. https://doi.org/10.1007/s41870-025-02723-6
[10] Sahoo, D., Liu, C., & Hoi, S. C. H. (2017). Malicious URL detection using machine learning: A survey. arXiv preprint arXiv:1701.07179.
[11] Phishing URL detection and interpretability with machine learning: A cross-dataset approach. (2026). Security and Privacy, 9(1), e70175. https://doi.org/10.1002/spy2.70175
[12] Chakravorty, A., Price, M., Elsayed, N., & ElSayed, Z. (2026). Context-aware phishing email detection using machine learning and NLP. arXiv preprint arXiv:2603.27326.
[13] Wiechmann, M. (2021). enron_spam_data: The Enron-Spam dataset preprocessed in a single, clean CSV file [Data set]. GitHub. https://github.com/MWiechmann/enron_spam_data
[14] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., et al. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825-2830.
[15] Salton, G., & Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. Information Processing & Management, 24(5), 513-523.
[16] Cormack, G. V. (2008). Email spam filtering: A systematic review. Foundations and Trends in Information Retrieval, 1(4), 335-455. https://doi.org/10.1561/1500000006
[17] Guzella, T. S., & Caminhas, W. M. (2009). A review of machine learning approaches to spam filtering. Expert Systems with Applications, 36(7), 10206-10222. https://doi.org/10.1016/j.eswa.2009.02.037
[18] Blanzieri, E., & Bryl, A. (2008). A survey of learning-based techniques of email spam filtering. Artificial Intelligence Review, 29(1), 63-92. https://doi.org/10.1007/s10462-009-9109-6
[19] Dada, E. G., Bassi, J. S., Chiroma, H., Abdulhamid, S. M., Adetunmbi, A. O., & Ajibuwa, O. E. (2019). Machine learning for email spam filtering: review, approaches and open research problems. Heliyon, 5(6), e01802. https://doi.org/10.1016/j.heliyon.2019.e01802
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Nagham Kamil Hadi, Aymen Adil

This work is licensed under a Creative Commons Attribution 4.0 International License.


