Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 9 (2023), Issue 2

Automatic Depression Level Detection

Authors

author email, Vedant Bhamre, Danish Tamboli, Trushant Jadhav, Pranav Kurle

Abstract

According to physiological research, there are a variety of variances in both speech and face movements. These facial and vocal expressions are shared by healthy and depressed people. On the basis of this information, we offer the Multimodal Attention Feature Fusion and a novel Spatio-Temporal Attention (STA) network technique that are utilized to get the multimodal representation of depression signals to be able to predict the amount of personal depression. Correctly, we first separate segmenting the speech amplitude spectrum and video into predetermined lengths and submitting them to the STA network, which focuses on the audio and video frames used to detect depression in addition to integrating the attentional processing of spatial and temporal information mechanisms. The output of the STA network's final full connection layer is where the audio and video segment-level functionality is acquired. In order to collect the changes in every aspect of the audio and segment-level features for videos and summarize them as an audio and video feature level, this study also provides the eigen evolution pooling approach. The MAFF is then used to create a multimodal representation composed of modal complementary data, which is then inputted into a support vector regression predictor to determine the severity of the depression. The utility of our strategy is illustrated by experimental findings on the depression databases for AVEC2013 and AVEC2014

Pages: 268 - 272