사업성과 BK21 FOUR 산업혁신 애널리틱스 교육연구단

논문

2025 Fashion Image Retrieval With Vision–Language Model Guided Fine-Grained Textual Attributes and Cross-Domain Contrastive Optimization

페이지 정보

작성자 관리자 작성일 26-07-15 09:36

본문

Author
Eojin Kim, Sangyeop Kim, Cholhwan Jung, Youngseok Hahm, Sungzoon Cho
Journal
IEEE Access
Vol
13
Page
192403-192415
Year
2025

Abstract

Fashion image retrieval (FIR) enables consumers to discover products using visual queries, however current methods struggle with fine-grained fashion differences and cross-domain gaps between user photos and professional catalogs. We propose a novel two-stage framework that combines vision-language model guided fine-grained textual attributes with weighted contrastive optimization to address these fundamental limitations. The first stage employs Sigmoid Loss for Language Image Pre-training (SigLIP) with systematic twelve-attribute fashion annotations using binary classification to accommodate multi-positive scenarios where single garments can be validly described by multiple attributes. This approach naturally handles the semantic richness of fashion items, where multiple descriptions may accurately characterize individual garments. The second stage introduces Weighted Multi Normalized Temperature-Scaled Cross-Entropy (NT-Xent) loss that strategically prioritizes street-to-shop matching while eliminating complex negative sampling requirements, enabling effective cross-domain learning between casual user photography and professional catalog imagery. Extensive experiments on Street2Shop and DeepFashion datasets demonstrate that our framework outperforms existing methods in terms of retrieval accuracy and zero-shot classification performance. Comprehensive ablation studies confirm the effectiveness of each proposed component in capturing detailed fashion characteristics and enabling practical cross-domain matching. This paper contributes a comprehensive framework addressing fashion retrieval challenges, novel application of SigLIP with specialized contrastive learning, and practical improvements for real-world applications.