GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Multimodal Cyberbullying Detection using Text, Emojis, and Images
Authors
Pulivarthi Sandhya, Atluri Aasritha, Golve Anusha, Pannala Renuka, D. Rajeswara Rao
Abstract
Cyberbullying has progressed from simple text and is now often shown through images, screenshots, and emojis, complicating detection. Conventional text-based models do not effectively harness these multimodal signals, leading to diminished precision in practical situations. To overcome this limitation, this paper proposes a multimodal cyberbullying detection system that combines Optical Character Recognition (OCR), deep learning text analysis, and visual feature extraction. The system utilizes OCR to extract text from images and processes it with DeBERTa embeddings paired with LSTM to capture contextual and sequential patterns. Moreover, emoji and visual indicators are examined through a CLIP-based feature extractor. A fusion approach integrates text and visual representations to categorize the input as toxic or non-toxic. The model attains impressive results with a validation accuracy of 94%, showcasing its capability in identifying cyberbullying through various modalities. The system is implemented through a Gradio-powered web interface, allowing for immediate evaluation of text and visuals.
Pages:
5760 - 5767