GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Deep NLP Models for Extracting Mentions of Low- Frequency Diseases from Medical Texts
Authors
Venkatesh N, Balajee Maram, Madichetty Mounika, P Archana, Sathishkumar V E
Abstract
Clinical texts and documents identifying mentions of diseases alongside monitoring and clinical support systems for rare diseases may be able to provide support for one another, though this synergy appears to be poorly explored, to our knowledge. The lack of NLP system configuration to the unique data set is the most plausible explanation for low performance in Sparse and Imbalance Data Frameworks. This research is about a deep learning hybrid architecture for NLP on BioBERT which uses contextual augmentation for extraction performance boosting and the model was tested on 50,000 clinically annotated documents. This model attained 91.2% and 83.7% on the F1 metric for prevalent and infrequent diseases, respectively, a 12.4% relative improvement over baseline models. The testimony of a rare disease expert was an important factor affecting the model's recall on infrequent diseases which increased by 18.5% also improving the overall F1 score.
Pages:
1198 - 1204