Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

FashionIQ-X: Multimodal Fashion Retrieval and Recommendation using Vision-Language Models

Authors

Viomesh Singh, Vira Shah, Vishal Arkalwar, Aary Tadwalkar, Utkarsh Tripathi Avni Raich

Abstract

This paper presents FashionIQ-X, a multimodal frame-work for semantic fashion retrieval and personalized recommendation using vision-language models. The pro-posed system combines BLIP-2-based caption genera-tion, CLIP embeddings, and FAISS vector indexing to enable natural-language search over large fashion cata-logues. In order to achieve higher accuracy in search and retrieval, captioning is performed by combining information from both the datasets and the generated captions, thus obtaining better semantic representation. Moreover, large language models can be leveraged to convert structural data about users into descriptive fashion query requests. Results on a dataset containing 20,491 fashion images show superior accuracy with regard to traditional approaches utilizing only metada-ta, scoring an accuracy of 82.4% in Recall@20. It is concluded that vision-language captioning and vector semantics retrieval is a viable approach to intelligent fashion recommendation.