2025 Fashion Image Retrieval With Vision–Language Model Guided Fine-Grained Textual Attributes and Cross-Domain Contrastive Optimization
페이지 정보
작성자 관리자 작성일 26-07-15 09:36본문
- Journal
- IEEE Access
- Vol
- 13
- Page
- 192403-192415
- Year
- 2025
Abstract
Fashion image retrieval (FIR) enables consumers to discover products using visual queries, however current methods struggle with fine-grained fashion differences and cross-domain gaps between user photos and professional catalogs. We propose a novel two-stage framework that combines vision-language model guided fine-grained textual attributes with weighted contrastive optimization to address these fundamental limitations. The first stage employs Sigmoid Loss for Language Image Pre-training (SigLIP) with systematic twelve-attribute fashion annotations using binary classification to accommodate multi-positive scenarios where single garments can be validly described by multiple attributes. This approach naturally handles the semantic richness of fashion items, where multiple descriptions may accurately characterize individual garments. The second stage introduces Weighted Multi Normalized Temperature-Scaled Cross-Entropy (NT-Xent) loss that strategically prioritizes street-to-shop matching while eliminating complex negative sampling requirements, enabling effective cross-domain learning between casual user photography and professional catalog imagery. Extensive experiments on Street2Shop and DeepFashion datasets demonstrate that our framework outperforms existing methods in terms of retrieval accuracy and zero-shot classification performance. Comprehensive ablation studies confirm the effectiveness of each proposed component in capturing detailed fashion characteristics and enabling practical cross-domain matching. This paper contributes a comprehensive framework addressing fashion retrieval challenges, novel application of SigLIP with specialized contrastive learning, and practical improvements for real-world applications.