GRENZE International Journal of Engineering and Technology
Vol. 9
(2023), Issue 1
An Overview of Speaker Recognition: Conceptual Framework and CNN based Identification Technique
Authors
Avinash Dhole, Vijaylaxmi Kadroli
Abstract
One of the major characteristics of an individual is to recognize a human voice. He is capable to do so, via digital devices or over the telephone. Acquiring this human trait, multiple technologies based on voice recognition have been developed to fulfil the purpose of biometrics and authentication. One such speech analysing technology is automatic speaker recognition (ASR) method that has been introduced to extract characteristics from the speaker’s voice and identify him as a genuine source of input. The primary aim of speaker recognition is to identify and verify a person using audio signals. Hence, this has become a dominating field of research in the domain of biometrics. However, several deep learning approaches based on convolutional neural networks have been adopted to enhance the overall system of biometrics. Therefore, in this paper we review feature components on speaker recognition using CNN and feature extraction methods such as MFCC. The major advantage of using this method over traditional identification is its representation ability to extract feature inputs including audio signals and networking structures. Further, we briefly describe all the main pieces of ASR related methodologies followed by evaluation metrics to enhance the overall recognition of the system. Finally, a few relatable challenges and future expansion of speaker recognition are mentioned at the closure of this review.
Pages:
2901 - 2908