GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
CurriculumLip: Progressive Visual Speech Recognition via Self-Supervised and Supervised Curriculum Learning
Authors
Harshita, SangeetaKumari, Aseem Madaan
Abstract
Despite advancements in deep learning, visual speech recognition (lipreading) is still challenging because of viseme confusion, variability in how speakers articulate themselves and the tiny amount of labelled training data available. The methodologies used to date typical-ly rely either on fully supervised training or on disjoint pretrial-fine-tune training pipe-lines, and therefore ignore progression of difficulty across structured utterances. This paper presents a novel three-stage curriculum-learning curriculum (CurriculumLip), which progressively incorporates supervised CTC training with each of the two visual self-supervised pretext tasks that comprise the curriculum: predicting temporal ordering of lip triplets (predicted in the future) and past predictive future features. The architecture employs a 3D CNN spatiotemporal front-end, followed by a multi-resolution Temporal Convolutional Network (TCN) encoder, and captures lip motion over local, medium, and long temporal lengths through parallel dilated convolution processes. CurriculumLip was evaluated on the GRID corpus with a word accuracy (WA) of 83.7%—an absolute difference of 5.5%—and an additional relative reduction of 35.9% in WER provided labels for the purpose of generating out-of-distribution sequence length modelling. Additionally, ablation studies support the contribution of the multiple stages of the curriculum as well as other tests that highlight reliability, robustness and ability to generalize (7.9% vs 15.6% accuracy loss) using different amounts of frames available at dropouts.
Pages:
1516 - 1524