GRENZE International Journal of Engineering and Technology
Vol. 11
(2025), Issue 2
Designing and Implementing AI Image Caption Bot with Speech Convertor
Authors
Ujwalla Gawande, Saloni Bhandirge, Prajakta Khandare, Shejal Hajaree, Vinita Kankate
Abstract
This research explores the development of an AI-powered system aimed at assisting visually impaired people in converting visual information into descriptive audio outputs. It makes use of advanced deep learning techniques to analyze and interpret image content, thereby generating accurate captions that are later converted into speech. A custom-built application was developed to make the system accessible, and users could upload images and receive descriptive feedback in text and audio. The model was trained on a diverse dataset and reached an accuracy of 80% for relevant captions. This solution is meant to enhance the independence of visually impaired users, providing a seamless way of interacting with visual content in order to showcase the full potential of AI in furthering accessibility and inclusion. By translating complex visual data into descriptive audio narratives, this system bridges the gap between visual content and auditory comprehension. The project further integrates advanced deep learning methodologies utilizing a CNN-LSTM for image captioning and integrating a state-of- the-art TTS engine to present natural audio feedback. The model was trained on an enriched dataset, including the Flickr8K corpus, allowing it to achieve 80% accuracy in generating relevant, context-aware captions. The app is implemented using Streamlit, allow ing users to input images and receive both textual and audio feedback instantly.
Pages:
2179 - 2183