GRENZE International Journal of Engineering and Technology
Vol. 8
(2022), Issue 1
Accuracy Comparison: An Approach of Unsupervised Machine Learning Algorithm for Assamese Speech Recognition
Authors
Devi Mampi, Sarma Manoj Kumar
Abstract
In the recent years the Automatic speech recognition systems have made vivid acts. Still, the key to making recognition more robust is to reduce the difference between training and testing with speech that commonly held. The postulation of ASR (Automatic speech recognition) is becoming less and less feasible, because ASR applications change from firmly measured to more usual situations with a erratic number of random sound bases. Decoding the speech source of interest while listening to several sound sources at the same time seems a more accurate description of the ASR process with different features that suits these challenging environments. The aim is to compare the accuracy with different features at different instance for robust ASR into two sub-problems:(a) identification/separation of the speech and noise using speech properties alone and (b)recognition based on the comparison of the accuracy and entropy. The basic assumption is that some regions of the speech time-frequency representation remain relatively unaffected by the noise, that they can be identified and that they alone are sufficient for ASR. In contrast to conventional techniques which require models of all sources in the auditory scene and their subsequent decoding even when only one of the sources is of interest, the techniques described in this paper make no such requirement. However, they are flexible enough to use this information if itis available. In the adaptation of a conventional Hidden Markov model (HMM) based ASR system two techniques are used for partial indication: (i) down grading of the state dispersals, so that only the likelihood of the reliable sections are evaluated; and (ii)attribution of the undependable sections by replacing the unreliable features with a single point from the community restricted deliveries. In the experiments, the reliable features are identified via clustering accuracy derived through Kmeans clustering. The potential of the techniques is indicated by using the clean speech to identify the reliable regions in the noisy speech.
Pages:
291 - 298