Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 10 (2024), Issue 1

Harmonizing Vision and Voice: A Review of Contemporary Research in Image Caption Generation and Text-to-Speech Synthesis

Authors

Yash Anand, Rudra Rawat, Aditya Sahu, Aayush Tandon, Vaibhav E. Narawade

Abstract

Automatically describing the content of a picture with meaningful and contextually appropriate textual descriptions is a difficult challenge at the interface of computer vision and natural language processing. This review paper provides a thorough summary of recent developments in picture caption generating methods, including both conventional strategies and cutting-edge deep learning-based solutions. We examine the changes in datasets, evaluation criteria, and model designs that have influenced this field's advancement. In order to improve the standard and variety of generated captions, we also integrate visual attention processes, transformer-based models, and reinforcement learning techniques. The focus of this study is on the algorithmic overlap between text to speech and visual

Pages: 2397 - 2401