GRENZE International Journal of Engineering and Technology
Vol. 10
(2024), Issue 1
Enhancing Web Security: An Efficient URL Phishing Classifier based on Deep Learning
Authors
Abhijna S, Vinay Prasad M S
Abstract
To procure sensitive data such as account Identification details, usernames, passwords, etc Attackers often develop new strategies, such as phishing, to fool users by utilizing fake websites. This paper presents and illustrates various methods for addressing phishing detection. To enhance user security and privacy, a web application is developed to determine if the URLs are malicious or benign. Malicious website detection can significantly decrease the probability of phishing attempts. In this paper, two distinct feature extraction techniques, Natural Language Processing (NLP) and TF-IDF vectorization, are employed to capture the intricate patterns within URLs. Evaluation and training are carried out by feeding the NLP features into Convolutional Neural Networks (CNN), Multilayer Perceptron’s (MLP), Long Short-Term Memory Networks (LSTM), and Deep Neural Networks (DNN). Initial experimentation with NLP features yields promising results, achieving up to 98.78% accuracy with CNN and 95.32% with LSTM. An accuracy of 94.03% and 94.64% is obtained by exploring TF-IDF vectorization using CNN and MLP models respectively. Precision, F1-score, recall, ROC and various metrics for evaluation provide a holistic assessment of the model's performance. Oversampling and class weight adjustment techniques address the implications of class imbalance. To enhance model generalization, Hyperparameter tuning, cross validation, and ensemble methods are emphasized. Overfitting concerns are mitigated using regularization techniques such as dropout. This approach involves Principal component analysis (PCA), a feature reduction technique to lower the number of features present
Pages:
337 - 345