Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 10 (2024), Issue 2

Enhancing K-Nearest Neighbors with Dimensionality Reduction: An Impact on Recognition Time and Accuracy

Authors

Priyadarshan Dhabe, Harshita Bhagat, Parth More, Vaishali More, Chinmay Saraf

Abstract

KNN [1] is the simplest and popular classifier used for pattern recognition tasks and predictions across various application domains. However, certain limitations hinder its efficiency. These include the higher computational cost due to excessive recognition time and more memory space. KNN requires significant computer memory as the complete training dataset needs to be stored in the computer memory. Thus, it is less suitable for larger training and testing datasets with higher dimensionality. KNN requires time and memory proportional to the size of training and testing datasets and their dimensionality. To reduce this handicap of KNN, in this paper, we propose a simple statistical method to reduce the dimensionality of features used in training and testing datasets. The proposed approach helps in reducing recognition time and memory space both. We did experimentation by using 3 datasets available on Kaggle [2] namely the breast cancer dataset [3], Wine quality dataset [4] and the diabetes dataset [5]. As per our experimentation, we observed, on an average, 27% reduction in recognition time as compared to using all the pattern features. We also found 18% to 30% reduction in number of features, without reduction in accuracy. In fact, for breast cancer and wine quality dataset, there is a slight increase in recognition accuracy too. Thus, we recommend this approach to use KNN for larger datasets with high dimensionality also.

Pages: 2942 - 2949