GRENZE International Journal of Engineering and Technology
Vol. 8
(2022), Issue 2
MarathiSarc: A Marathi Tweets Dataset for Automatic Sarcasm Detection of Marathi Tweets
Authors
Pravin K. Patil, S. R. Kolhe
Abstract
Sarcasm Detection is a task of predicting whether the given text is sarcastic or not. Considering the challenges in detecting sarcasm in a sentiment bearing text, sarcasm detection has become one of the hot research areas in Natural Language Processing. Considerable amount of research has been done in this area for foreign languages such as English, Czech, Italian, Dutch, and Indonesian etc. Small amount of work in this area is also available for Indian languages such as Hindi, Tamil, and Bengali etc. However, Marathi being the third most popular language in India lags far behind in this area. One of the most crucial reasons for this is the absence of proper dataset. In this paper, we present MarathiSarc - a dataset of labeled Marathi tweets for sarcasm detection. This dataset contains 2400 tweets in Marathi language. Further we also discuss our dataset collection and the annotation policy. Finally we present some of the baseline sarcasm detection experiments performed on this dataset. The dataset is made publicly available at https://github.com/patilpravin91/MarathiSarc.
Pages:
108 - 114