Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 1

Hybrid Feature Selection of Electronic Health Records for Heart Disease Prediction using Natural Language Processing Techniques

Authors

V. Sravanthi, Sheshikala Martha

Abstract

Cardiovascular Disease is a prominent cause of death in nearly all the populations globally and early detection of the risk factors is of utmost importance in facilitating cardiovascular prevention and treatment. In this research, a complete methodology to extract and use clinically meaningful features out of unstructured Electronic Health Records (EHRs) was presented to aid in the prediction of heart disease. The first step involved the use of Natural Language Processing (NLP) based techniques to extract useful clinical information out of freetext narratives such as discharge summaries, progress notes and laboratory reports. Entity linking along with clinical concept extraction was then carried out based on domain specific ontologies (e.g., SNOMED CT, UMLS) to normalize extracted entities and disambiguate them. This guaranteed the representation of semantically similar terms that would be used in further analysis. In the second step, a hybrid feature selection model was developed and applied to extract important predictors of heart disease among the structured and extracted unstructured data. This framework combined statistical (filter-based), heuristic (wrapper-based) and deep learning-based (embedded) approaches. Such methods like chi-square, mutual information, recursive feature elimination and L1-regularized models were integrated in a systematic way. Secondly, a domain expert knowledge in the field of cardiology was integrated to filter the set of extracted features and select only those that are clinically relevant. This method shows potential in achieving improved clinical decision support through the utilization of the entire range of data incorporated in EHRs.