GRENZE International Journal of Engineering and Technology
Vol. 11
(2025), Issue 2
Multimodal Emotion Recognition Integrating Facial Expressions and ECG Signals with Deep Learning Model
Authors
Venkatesh Koreddi, Jyothi Chinta, Hima Padala, Praveen Kumar Kovvuru
Abstract
Understanding human emotions using multiple data points is an expanding research area with significant applications in mental health. In this study, we present a novel deep learning model that integrates visual and physiological information to decode and classify emotions with higher accuracy compared to previous approaches that rely on a single data type. By leveraging Mobilenetv2 for efficient real-time facial feature recognition and a specialized convolutional neu- ral network (CNN) for capturing intricate patterns in electrocardiogram (ECG) signals, our model effectively captures both external and internal emotional cues. The data are independently processed to extract robust features, which are then combined using an innovative fusion technique with an attention mechanism to optimize the contribution of each data type. The fusion of these two modalities led to a significant improvement in emotion prediction accuracy, achieving 94.3%. While MobileNetV2 excelled at recognizing emotions such as “Happy†and “Neutral†from facial expressions, and the CNN model effectively classified physiological states like “Happy†and “Sad,†the integrated model outperformed individual models, particularly for harder-to-distinguish emotions such as “Fear†and “Disgust.†This multimodal approach surpasses single-modality methods, offering a scalable, robust framework suitable for various applications, including user-friendly learning systems and emotion recognition devices. This research advances the understanding of complex emotions and sets a new benchmark in emotion recognition by emphasizing the critical role of integrating visual and physiological data.
Pages:
678 - 683