GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Marathi Paraphrase Detection using Lexical Similarity and MahaBERT - based Semantic Analysis
Authors
Sheetal R. Dehare, Aarti P. Raut, Rajashri G. Kanke, C. Namrata Mahender
Abstract
Paraphrase detection is an essential problem in Natural Language Processing (NLP) concerned with identifying if two sentences have similar meanings, but are said in different ways or presented in different structures. Although tremendous strides have been made for languages with abundant resources like English, research in Indian languages like Marathi is still limited, as there are few annotated datasets and language-specific models available. In this study, a Marathi paraphrase detection system has been created with both lexical and deep learning techniques. A custom dataset of 3,540 sentences was developed from the textbooks from Balbharati, with each sentence having three paraphrased variants created manually. The system compares two approaches: lexical-based Jaccard similarity and contextual semanticbased cosine similarity using a transformer-based approach that captures contextual semantic representations, MahaBERT. Experimental results show that Jaccard similarity would only compare on the surface level and not recognize paraphrases when different words were used. On the other hand, MahaBERT demonstrates a remarkable ability to represent semantic links between sentences and substantially outperforms the lexical approach. The accuracy of this model is 88%, with a precision of 87%, a recall of 89% and F1 of 88%. The results highlight the importance of deep learning-based contextual models for paraphrase detection in morphologically rich and low-resource languages like Marathi. This work aims at promoting Marathi NLP and act as a stepping stone for further research on semantic similarity and other applications.
Pages:
6065 - 6072