GRENZE International Journal of Engineering and Technology
Vol. 11
(2025), Issue 2
Hybrid Retrieval Systems: Combining Graph-Structured and Embedding-based Models for Efficient Document Processing
Authors
Jeevanantham G, Sanmuga Priya M, Barakkath Nisha U
Abstract
In the current realm of information retrieval, the adoption of hybrid machine learning models offers a significant opportunity to enhance document processing and knowledge-driven question answering. This paper presents the development of a Hybrid Retrieval System that effectively merges graph-based and vector-based retrieval methodologies. The system leverages HuggingFace embeddings alongside Neo4j graph databases to improve the precision and relevance of information extracted from large document collections. A key element of our strategy involves constructing both a knowledge graph and a vector store from unstructured documents, employing semantic chunking to break down text into significant segments. We propose a robust retrieval framework that combines vector search with knowledge graph navigation, thereby optimizing the retrieval of contextually pertinent answers. Additionally, the integration of Contextual Knowledge Distillation (CKD) facilitates the extraction and compression of relevant subgraphs in response to user queries, thereby enhancing overall retrieval accuracy. The system's capabilities are demonstrated through various use cases, highlighting improvements in response accuracy, document indexing efficiency, and the speed of knowledge retrieval. This hybrid methodology not only streamlines the retrieval process but also provides scalability, flexibility, and adaptability across diverse sectors, including legal research and healthcare. The results emphasize the potential of the Hybrid Retrieval System to transform the construction, querying, and distillation of large-scale knowledge bases tailored to specific domain requirements.
Pages:
15185 - 15189