Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

ICIS: A Vision - Language Dataset of Indian Streets for Culturally Aware Image Captioning and Benchmarking

Authors

J Bhuvana, Sandhya Giribabu, Shinigdapriya Sathish, Vidarshanaa Saravanavel

Abstract

Existing large-scale Vision-Language (V+L) datasets are predominantly focused on Western urban environments, resulting in a geographical bias that hinders model performance in visually complex and culturally diverse settings like urban India. To bridge this gap, we introduce Images and its Captions of Indian Streets (ICIS), a novel dataset comprising 1,350 high-resolution images capturing heterogeneous traffic, crowded marketplaces, and traditional attire. An analysis of 5,400 captions reveals a remarkably low inter-annotator agreement score, which validates the high subjectivity and diversity of human interpretation inherent in these scenes. Benchmarking with the state-of-the-art BLIP model demonstrates a significant performance gap; while semantic metrics suggest general thematic understanding, lexical overlap and multimodal alignment scores remain notably low. Experimental evaluation using BLIP, BLIP-2, and GIT observes high semantic similarity (BERTScore ? 0.89) but extremely low lexical overlap (BLEU-4 < 0.06), showing the linguistic diversity and visual complexity of Indian street scenes. These results indicate that current V+L models struggle with the culturally nuanced descriptions required for unstructured environments. ICIS provides a challenging benchmark to advance the development of globally robust and culturally aware computer vision systems for applications in smart cities and autonomous navigation.

Pages: 129 - 136