Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

An Integrated OCR–NLP Framework for Printed and Handwritten Text Recognition

Authors

Ravi P, Y. H. Sharath Kumar, Thejashwini M A, Thanushree S R, Sonashree M.S, Nuthan Mourya N

Abstract

Optical Character Recognition (OCR) systems often encounter significant challenges when converting handwritten and degraded texts into digital format, primarily due to noise, inconsistent writing styles, and low-quality inputs. To overcome these issues, this research work introduces an integrated OCR–NLP framework that integrates OCR processing model and language model post-processing and correction. This proposed system applied DenseNet-121 for text-type classification, Tesseract OCR for printed text extraction, and OCR Space for handwritten text extraction. This extracted text is further refined by using PHI-3 language method for grammar correction, structural and semantic inconsistencies. Experimental result analysis on a mixed dataset of printed text and handwritten demonstrate an 85% reduction in grammatical errors and enhanced text readability from 69.2 to 91.4 on the Flesch scale. These results confirm that the proposed integrated OCR–NLP framework significantly improves recognition accuracy and produces contextually accurate output suitable for real-world text digitization applications.