Deep Metric Learning with MobileNetV2 and Triplet Loss for Efficient Image Retrieval

Authors

  • Hersh M. Hama Software Engineering Department, College of Engineering, University of Raparin, Ranya, Sulaymaniyah, Iraq. Author https://orcid.org/0000-0003-1077-0933
  • Kamaran H. Manguri Software Engineering Department, College of Engineering, University of Raparin, Ranya, Sulaymaniyah, Iraq. Author https://orcid.org/0000-0001-8567-3367
  • Saman M. Omer Software Engineering Department, College of Engineering, University of Raparin, Ranya, Sulaymaniyah, Iraq. Author https://orcid.org/0000-0002-5517-9066

DOI:

https://doi.org/10.24017/science.2026.2.4

Keywords:

Content-based image retrieval, MobileNetV2, Deep metric learning, Dimensional feature reduction

Abstract

In the context of today's digital system applications, content-based image retrieval (CBIR) is an essential tool for managing large visual repositories. Current CBIR systems, however, face a trade-off between retrieval accuracy and computational efficiency, caused by high-dimensional representations and their heavy memory and query-time requirements. To address this problem, this paper proposes an efficient CBIR system   based on deep metric learning. The system integrates the MobileNetV2 network for feature extraction, a trainable linear projection head, and triplet-loss optimization. The aim of this integration is to learn compact, discriminative embeddings that improve retrieval accuracy while maintaining a fast search time. The system was trained and tested with the Corel-1K database, and tested with 200 query images. The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512. This reduction not only improves retrieval accuracy, but also reduces query latency by more than 59%. The proposed model achieved a precision at 10 (P@10) of 97.65% and an average precision at 10 (AP@10) of 97.11%. Moreover, the framework maintained performance at greater retrieval depth, achieving P@20 and AP@20 scores of 97.30% and 96.74%, respectively, exceeding those of several more complex methods. The learned embeddings are also robust across image classes, as further confirmed by the qualitative and class-wise analyses. Overall, the results confirm that compact metric-learning representations provide a practical, scalable, and efficient approach for contemporary CBIR applications.

Downloads

Published

31-08-2026

Issue

Section

Pure and Applied Science

Similar Articles

51-60 of 173

You may also start an advanced similarity search for this article.