Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 7 (2021), Issue 2

Impact of Preprocessing and Feature Representation in the Classification of Documents with Abusive Content

Authors

Divya Ann Kurien, Namitha S, Vinaya Elma Givson, Ansamma John

Abstract

The use of social media has been increasing over the past decade and has skyrocketed last year due to the pandemic. With the increase in engagement, the amount of abusive text generated has also increased massively, that manually identifying and removing it has become a difficult undertaking. Though automated systems are built, its performance in categorizing text into abusive is highly affected by noise and outliers in unstructured text from social media. The performance of abusive language systems can be improved, by the appropriate structured representation of input text documents with suitable preprocessing stages. The selection of preprocessing steps to improve the overall efficiency of abusive text detection systems depends on the nature of words, language based features and abbreviations which are specific to the dataset. This work focuses on illustrating the appropriateness of different preprocessing steps and its individual or combinational selection based on the features of the dataset used, by transforming the document into a structured form with bag of words or term frequency-inverse document frequency representation methods. It is observed that by combining various preprocessing methods, overall computational efficiency is improved.

Pages: 77 - 86