GRENZE International Journal of Engineering and Technology
Vol. 8
(2022), Issue 1
A Comparative Study of Re-Sampling Techniques on Machine Learning Model Performance
Authors
Pavitha N, Ashwini B P, Savithramma R M, R Sumathi
Abstract
In this data era, gaining insights into available data through machine learning to provide business solutions is the trend. Data in few domains such as banking, disease diagnosis and fraud detection, etc, are imbalanced, and to apply machine learning on these require overcoming the imbalance. Applying re-sampling techniques is one the most followed method to overcome this. Various re-sampling methods are available with the effort of researchers across the world. In this context, an attempt has been made in this article to compare various resampling techniques on an imbalanced dataset; a Portuguese bank marketing dataset from UCI Machine learning repository. The impact of various re-sampling techniques on the performance of a classic machine learning model; Logistic regression is analyzed. Binary classification is conducted in-order to forecast the subscription status of a customer for the term deposit. The model is evaluated for various performance evaluation metrics such as AUC-ROC (Area Under Curve-Receiver Operating Characteristic), Accuracy, Precision, Recall and F1-score. The obtained empirical results showed that the K-Means SMOTE is the best (94% accuracy), whereas, SMOTE+ENN is the second best (92% accuracy) suitable model for logistic regression while using the Portuguese banking dataset.
Pages:
227 - 234