Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 9 (2023), Issue 1

Automated Text Extraction from Images using Optical Character Recognition

Authors

Arya Karambelkar, Param Mamania, Niti Shah, Nandana Prabhu

Abstract

In the modern day, practically everything is mechanized, and data is saved and shared digitally. In other circumstances, though, the data won't be digitized, and it could be necessary to separate the text from it to keep it digitally. The technique of extracting text using optical character recognition has been entirely transformed by the newest technologies, such as text recognition software. So, the goal of the software is to automatically retrieve text from images. This requires an optical character recognition system that incorporates multiple algorithms. Tesseract is currently the most accurate optical character recognition engine developed by HP Labs and currently owned by Google. This paper talks about extraction of text from images using text localization, segmentation, and binarization techniques.

Pages: 270 - 275