Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

Video Captioning with Text and Audio Output

Authors

Sheela N, VaniAshok, Manimala S, Chinmay G S

Abstract

Complex dynamics are inherent in real-world videos and the generation of opendomain video description should be capable of handling temporal structure and allow the input and output of the video description to be of arbitrary length. For this solution, we used an endto- end sequence to sequence model for generating captions for videos. Here, we utilize Recurrent Neural Networks (RNN) especially Long Short-Term Memory (LSTM) as they are well suited for image caption generation and are among the best performing models today. In order to provide a summary of the incident in the video clip, our suggested LSTM model maps a sequence of video frames with a sequence of words using pairs of textual and video descriptions as input. By construction, our Language Model (LM) assumes both the temporal structure of the frame sequence and the sequence model of the generated sentences. We evaluate multiple iterations of the model using various video visual attributes on the Microsoft Research Video Description (MSVD) dataset in order to examine various aspects of the suggested model.