GRENZE International Journal of Engineering and Technology
Vol. 10
(2024), Issue 2
Integrating Tesseract OCR with Large Language Models for Enhancing Textual Understanding
Authors
A V Sriharsha, Mekala Bhavana, Syamala Tejaswee, Mynasaheb Bushra Ahmed, Peddi Reddi Jaipal Reddy
Abstract
Optical Character Recognition (OCR) is a technology that automatically extracts textual information from document images. An LLM, a large language model, is a neural network designed to understand, generate, and respond to human-like text. These models are deep neural networks trained on massive amounts of text data, sometimes encompassing large portions of the entire publicly available text on the internet. LLMs require textual data to be converted into numerical vectors, known as embeddings since they can't process raw text. Embeddings transform discrete data (like words or images) into continuous vector spaces, making them compatible with neural network operations. The combined strengths of Tesseract OCR's image-based text extraction and the deep language understanding of LLMs enhance the accuracy and depth of textual comprehension. The proposed integrated model establishes a novel paradigm in enhancing textual understanding, progress in document analysis, information retrieval, and various NLP applications.
Pages:
1626 - 1633