Aug 2026· Kurdistan Journal of Applied Research· 0 citations· 41 references
TL;DR
The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512, which improves retrieval accuracy, but also reduces query latency by more than 59%.
Abstract
In the context of today's digital system applications, content-based image retrieval (CBIR) is an essential tool for managing large visual repositories. Current CBIR systems, however, face a trade-off between retrieval accuracy and computational efficiency, caused by high-dimensional representations and their heavy memory and query-time requirements. To address this problem, this paper proposes an efficient CBIR system based on deep metric learning. The system integrates the MobileNetV2 network for feature extraction, a trainable linear projection head, and triplet-loss optimization. The aim of this integration is to learn compact, discriminative embeddings that improve retrieval accuracy while maintaining a fast search time. The system was trained and tested with the Corel-1K database, and tested with 200 query images. The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512. This reduction not only improves retrieval accuracy, but also reduces query latency by more than 59%. The proposed model achieved a precision at 10 (P@10) of 97.65% and an average precision at 10 (AP@10) of 97.11%. Moreover, the framework maintained performance at greater retrieval depth, achieving P@20 and AP@20 scores of 97.30% and 96.74%, respectively, exceeding those of several more complex methods. The learned embeddings are also robust across image classes, as further confirmed by the qualitative and class-wise analyses. Overall, the results confirm that compact metric-learning representations provide a practical, scalable, and efficient approach for contemporary CBIR applications.
Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this to...
The findings provide practical guidelines for selecting augmentation techniques that maximize test diversity while preserving realistic image characteristics, thereby enabling the construction of comprehensive and effective test suites for image retrieval systems while reducing the cost of manual data labeling through...
Yehan De Silva, Anirudh Sridhar, Armin Lotfy et al.· 0 citations
The rapid growth of the fashion industry has increased the demand for effective retrieval of visually similar clothing items from large and diverse image collections. This paper proposes a content-based image retrieval (CBIR) framework for fashion images that combines multiple feature extraction methods with scalable s...
Truc Nguyen, N. Nguyen, V. Tran et al.· International Conference on...· 0 citations
An Enhanced Content-Based Image Retrieval Using Fusion of Feature Representation with Optimised Image Similarity Measures (CBIRFR-OISM) approach is proposed to effectively enhance the system’s capability to retrieve visually and contextually similar images.
This work revisits the efficacy of simple linear interpolation within an embedding space, and introduces SRAIN, the first framework that dynamically predicts query-specific interpolation weights, and achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval.
Boseung Jeong, T. Park, Donghyeon Kwon et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.