GRENZE International Journal of Engineering and Technology
Vol. 11
(2025), Issue 2
Enhanced Image Captioning with Deep Learning: Leveraging Attention Mechanism for Improved Image Description
Authors
Pravin Patil, Bhupesh Suryawanshi, Om Nikhade, Nayan Magare, Kamlesh Mali
Abstract
With the unprecedented growth in visual content around the internet, it’s important to give images relevant captions. The automation of image captioning is beneficial across various areas starting from making images more accessible to blind people to making images easily searchable in social media and e-commerce sites. Over time, this technology has transitioned from basic, rule-driven methods to advanced deep learning techniques that can generate complex and precise descriptions. Previous methods, which were based on commanddesigned features, frequently produced ambiguous or general captions. A comparison of older technologies (including the one being referred here) vs. modern technology such as deep learning, which can generate meaningful reports about images. With the use of CNNs to learn visual features and RNNs to generate relevant and coherent sentences, this project tries to improve image captioning techniques. Attention mechanisms further improve the effectiveness of the task, allowing the attention system to target only on the meaningful parts of the image. Which results in less accurate (textually) but more visual context-based captions. This paper advances the field of image captioning by proposing a system that produces captions that closely aligned with human-generated captions. The outcomes of this project will improve user experience and serve as helpful solutions in these cases where visual data understanding is required.
Pages:
982 - 988