GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Vision based Lip-Reading System using Deep Learning
Authors
Ashwini G, KN Meghana, L Sanjana, Shwetha M, Darshan DM
Abstract
We present a deep-learning–based automated lip-reading system designed to recognize spoken words from silent video input. The system eliminates the need for audio by analyzing only the movement of the speaker’s lips. A hybrid architecture combining Convolutional Neural Networks (CNN), Long Short-Term Memory networks (LSTM), and Transformer encoders performs feature extraction, temporal modeling, and attention-based sequence decoding. The model is supported by a lightweight Flask-based user interface that enables real-time inference. Multiple security and reliability features—including dataset preprocessing, frame normalization, and model confidence filtering—ensure robustness under different lighting and speaking conditions. This system makes visual speech recognition more accessible, scalable, and usable in practical applications such as assistive communication, surveillance, and silent speech interfaces.
Pages:
3097 - 3101