GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Vision-Language Integration: Transformer-based Generative AI for Image Captioning to Improve Accessibility and Visual Content Understanding
Authors
N. Sandeep Chaitanya, M. Mohith, P. Vishnu Teja, P. V. Narendra Reddy, V. Varaprasad Reddy
Abstract
This project demonstrates a notion of a vision language corresponding generative artificial intelligence system in automated image captioning and improvement in accessibility. The suggested system comprises of a couple of transformers-based visual encoder and a generative language display, to interpret visual data and produce semantically significant natural language explanations. A trained BLIP (Bootstrapping Language-Image Pre-training) model is applied to infuse visual characteristics and correlate the visual characteristics to linguistic frames with the ultimate output being to produce the correct captions. This system facilitates both conditional and unconditional captioning in this manner, making it possible to specify the descriptions of a user as well as process an interpretation of pictures independently. The generated captions are further converted to speech to make it more inclusive with the help of a text-to-speech module, which is allowing visual impaired users to gain access to and interpret the visual information and is doing so via a different channel; audio feedback. A streamlit based interface is developed that offers interactive space of upload image, generation caption, language selection and the caption track. The architecture offers contextual knowledge, enhanced semantic knowledge and live usability. The suggested method demonstrates the effectiveness of transformer-based multimodal learning in bridging the sensory gap between visual perception and human language. To offer an adaptive and scalable and accessibility-oriented system to interpret automated images in various domains of applications, the system does the job of assistive technologies, intelligent content understanding, and human-computer interaction.
Pages:
162 - 170