Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 1

Multimodal Speech Recognition: Enhancing Real-Time Lip Reading with Audio-Visual Fusion

Authors

S.Sangeetha Mariammal, S. Saranya, K. Harsida, M. Rashmika, K.Shiyamala Devi

Abstract

Lip-reading systems, or Visual Speech Recognition (VSR), have become an essential area of research in the context of automatic speech recognition, particularly in noisy environments or when audio is unavailable. However, traditional visual-only lip-reading models face significant challenges due to viseme confusion, where similar-looking mouth shapes lead to difficulties in distinguishing between phonemes. Additional factors, such as variations in lighting, camera angle, and head pose, further complicate the accuracy of visual-based recognition. This paper presents a novel multimodal approach that combines visual lip movements with audio-based speech recognition to overcome these limitations. By leveraging deep learning techniques, including transformer models for visual speech recognition and audio processing, the proposed system enhances transcription accuracy in real-time, even under challenging conditions. The integration of both audio and visual inputs results in more robust performance, reducing errors caused by ambiguous visual cues. Experimental results demonstrate significant improvements in Word Error Rates (WER), showing that multimodal fusion can effectively address the inherent challenges of real-world lip-reading applications. The system’s potential for real-time deployment is validated through experiments on benchmark datasets Grammar-based Recognition of Isolated Digits (GRID) and Lip Reading Sentences 3 (LRS3), with promising results for applications in assistive communication, accessibility tools, and silent-speech interfaces.