GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Biometric Emotion Profiling Framework for Safe Audio Deep Fake Detection in Voice Authentication Systems
Authors
Preeti Shamrao Wathore, P R Murkute, S D Pingle
Abstract
Generative artificial intelligence is moving along quickly, making audio deepfakes much more realistic and easier to get to. This is a big problem for voice- based authentication systems. Traditional methods for finding audio deepfakes mostly use spectral or speakerdependent information. These methods often miss small emotional and behavioral anomalies that synthetic speech generation models add. To overcome this constraint, this study introduces a Biometric Emotion Profiling Framework for secure audio deepfake detection in voice authentication systems. The framework combines emotion-aware acoustic feature extraction (Mel- Frequency Cepstral Coefficients (MFCCs), pitch dynamics, spectral contrast, and signal energy) with biometric voice analysis and a Random Forest classifier to reliably tell the difference between real human speech and fake deepfake audio. The system is assessed with CREMA-D for authentic emotional speech and WaveFake for artificial deepfake audio. Experimental results show that the classification accuracy on a held-out test set is 100%. The precision, recall, and F1-score for both classes are 1.00, and five-fold cross- validation gives a mean accuracy of 1.0 with no volatility. An examination of feature importance suggests that the MFCC and pitch features that are based on emotions are the most important for telling deepfakes apart. The results show that adding biometric emotion profiling makes it much easier to find deepfakes and makes voice authentication systems safer.
Pages:
1570 - 1576