GRENZE International Journal of Engineering and Technology
Vol. 10
(2024), Issue 1
Harmonizing Vision and Voice: A Review of Contemporary Research in Image Caption Generation and Text-to-Speech Synthesis
Authors
Yash Anand, Rudra Rawat, Aditya Sahu, Aayush Tandon, Vaibhav E. Narawade
Abstract
Automatically describing the content of a picture with meaningful and contextually appropriate textual descriptions is a difficult challenge at the interface of computer vision and natural language processing. This review paper provides a thorough summary of recent developments in picture caption generating methods, including both conventional strategies and cutting-edge deep learning-based solutions. We examine the changes in datasets, evaluation criteria, and model designs that have influenced this field's advancement. In order to improve the standard and variety of generated captions, we also integrate visual attention processes, transformer-based models, and reinforcement learning techniques. The focus of this study is on the algorithmic overlap between text to speech and visual
Pages:
2397 - 2401