GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Early Identification of Students Requiring Academic Support: A Time-Sensitive Predictive Framework using Interpretable Machine Learning
Authors
Kuldeep Vayadande, Aditya Namdev Mali, Soham Milind Lonkar, Arjun Vishwas Mane, Sarthak Mahesh Patil, Parth Sagar Patil1
Abstract
At-risk students are becoming a significant problem that requires timely identification in order to provide assistance. The intervention is usually untimely, and this explains the effects associated with traditional models including poor academic results and increased number of dropouts. Conventional approaches to student evaluation typically involve exams conducted at the end of a semester and cannot provide timely assistance. Machine learning is employed in relation to existing methods, but is not capable of identifying at-risk students at the beginning of a semester because of insufficient amount of data, also known as the cold start problem. Moreover, numerous highly predictive black-box models with poor interpretation are not beneficial for teachers to understand risks associated with specific students. This project suggests developing an Early Warning Stu-dent Performance Prediction System to identify at-risk students in the first 6-8 weeks of a semester. The algorithms of machine learning utilized in the system include Extreme Gradient Boosting (XGBoost) and Random Forest. In order to ensure the lack of transparency of model predictions, Shapley Additive Explanations (SHAP) are utilized, thereby making the framework locally interpretable and helping educators find out individual reasons for predicting a risk in case of a particular student. With its locally interpretable AI and prediction based on the most primitive form of machine learning model, the proposed framework can be regarded as a means of conducting academic interventions based on information. A highly robust framework has been developed using XGBoost model with accuracy of 88.7 percent, precision of 87.4 percent and recall of 86.9 percent. The results suggest that the proposed methodology can reach a perfect balance between predictability and interpretability.
Pages:
5650 - 5657