Skip to content
Open access

Deep Metric Learning with MobileNetV2 and Triplet Loss for Efficient Image Retrieval

Aug 2026 · Kurdistan Journal of Applied Research · 0 citations · 41 references

TL;DR

The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512, which improves retrieval accuracy, but also reduces query latency by more than 59%.

Abstract

In the context of today's digital system applications, content-based image retrieval (CBIR) is an essential tool for managing large visual repositories. Current CBIR systems, however, face a trade-off between retrieval accuracy and computational efficiency, caused by high-dimensional representations and their heavy memory and query-time requirements. To address this problem, this paper proposes an efficient CBIR system   based on deep metric learning. The system integrates the MobileNetV2 network for feature extraction, a trainable linear projection head, and triplet-loss optimization. The aim of this integration is to learn compact, discriminative embeddings that improve retrieval accuracy while maintaining a fast search time. The system was trained and tested with the Corel-1K database, and tested with 200 query images. The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512. This reduction not only improves retrieval accuracy, but also reduces query latency by more than 59%. The proposed model achieved a precision at 10 (P@10) of 97.65% and an average precision at 10 (AP@10) of 97.11%. Moreover, the framework maintained performance at greater retrieval depth, achieving P@20 and AP@20 scores of 97.30% and 96.74%, respectively, exceeding those of several more complex methods. The learned embeddings are also robust across image classes, as further confirmed by the qualitative and class-wise analyses. Overall, the results confirm that compact metric-learning representations provide a practical, scalable, and efficient approach for contemporary CBIR applications.

Read PDF

Similar papers

Open access Sep 2026

Attention-enhanced vision transformer hashing for hybrid image retrieval

Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this to...

Uyen Nguyen, Hoai Ba, Quynh Dao Thi Thuy · 0 citations
Review Aug 2026

Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems

The findings provide practical guidelines for selecting augmentation techniques that maximize test diversity while preserving realistic image characteristics, thereby enabling the construction of comprehensive and effective test suites for image retrieval systems while reducing the cost of manual data labeling through...

Yehan De Silva, Anirudh Sridhar, Armin Lotfy et al. · 0 citations
Conference Sep 2026

VNIU-VNR50: a new dataset and framework for fashion image retrieval with vision transformer-based analysis

The rapid growth of the fashion industry has increased the demand for effective retrieval of visually similar clothing items from large and diverse image collections. This paper proposes a content-based image retrieval (CBIR) framework for fashion images that combines multiple feature extraction methods with scalable s...

Truc Nguyen, N. Nguyen, V. Tran et al. · 0 citations
#edge computing Open access Sep 2026

Fusion of local and global feature representation via optimised transfer learning approach on enhanced content-based image retrieval systems

An Enhanced Content-Based Image Retrieval Using Fusion of Feature Representation with Optimised Image Similarity Measures (CBIRFR-OISM) approach is proposed to effectively enhance the system’s capability to retrieve visually and contextually similar images.

Ravindar Karampuri, Sushama Rani Dutta · 0 citations
Preprint Aug 2026

Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

This work revisits the efficacy of simple linear interpolation within an embedding space, and introduces SRAIN, the first framework that dynamically predicts query-specific interpolation weights, and achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval.

Boseung Jeong, T. Park, Donghyeon Kwon et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.