VNIU-VNR50: a new dataset and framework for fashion image retrieval with vision transformer-based analysis
Abstract
The rapid growth of the fashion industry has increased the demand for effective retrieval of visually similar clothing items from large and diverse image collections. This paper proposes a content-based image retrieval (CBIR) framework for fashion images that combines multiple feature extraction methods with scalable similarity search. We also introduce VNIU-VNR50, a real-world dataset containing 1,311 images across 50 clothing categories, designed to reflect practical variations in lighting, viewpoints, and occlusions. The framework evaluates RGB Histogram, Local Binary Pattern (LBP), ResNet50, EfficientNetV2, and Vision Transformer (ViT), together with Facebook AI Similarity Search (Faiss) for efficient indexing and retrieval. Experimental results on both DeepFashion and VNIU-VNR50 show that deep learning-based methods outperform handcrafted features overall, with ViT achieving the best retrieval performance and reaching a mean Average Precision (mAP) of 0.791 on VNIU-VNR50. The results also reveal a trade-off between retrieval accuracy and computational efficiency. These findings demonstrate the effectiveness and practical value of the proposed framework for real-world fashion image retrieval. The dataset and source code are publicly available for research purposes.