GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
A Deep Learning–based Multimodal Cyberbullying Detection Framework using RoBERTa ResNeXt 50 and Cross-Modal Attention
Authors
Anurag Pandey, Jeba Nega Cheltha, Sonal Gandhi
Abstract
On social media Cyberbullying from last text form to fusion of text and images, sending message with hate and bullying. It is difficult to find this kind of abuse because it’s not about just a one thing it’s like a what a word or sentence says and what a image says so it’s trick to found that so, how they work together. In this paper we present a deep learning framework both of these two. It’s like a team RoBERTa is Text Expert, ResNeXt 50 is Image Expert and fusion is like a team manager who connect both of them and work together. You know how this thing works? It’s really clever! The attention mechanism is like a spotlight let the system zoom in on parts of a picture and match them to text. model catches hidden hate by understanding the context that other systems miss. Putted to test using both Bengali and Facebook meme datasets. It beat the existing technology, showing better accuracy and performance across the board. This research tooked a stand for fusion of pictures and text to detect cyberbullying. That proves that attention-based models is a super promising way to understand media without on the words.
Pages:
1795 - 1799