Deep Metric Learning with MobileNetV2 and Triplet Loss for Efficient Image Retrieval
DOI:
https://doi.org/10.24017/science.2026.2.4Keywords:
Content-based image retrieval, MobileNetV2, Deep metric learning, Dimensional feature reductionAbstract
In the context of today's digital system applications, content-based image retrieval (CBIR) is an essential tool for managing large visual repositories. Current CBIR systems, however, face a trade-off between retrieval accuracy and computational efficiency, caused by high-dimensional representations and their heavy memory and query-time requirements. To address this problem, this paper proposes an efficient CBIR system based on deep metric learning. The system integrates the MobileNetV2 network for feature extraction, a trainable linear projection head, and triplet-loss optimization. The aim of this integration is to learn compact, discriminative embeddings that improve retrieval accuracy while maintaining a fast search time. The system was trained and tested with the Corel-1K database, and tested with 200 query images. The results show that training the projection head with triplet loss improves embedding quality while reducing the embedding dimensionality from 1280 to 512. This reduction not only improves retrieval accuracy, but also reduces query latency by more than 59%. The proposed model achieved a precision at 10 (P@10) of 97.65% and an average precision at 10 (AP@10) of 97.11%. Moreover, the framework maintained performance at greater retrieval depth, achieving P@20 and AP@20 scores of 97.30% and 96.74%, respectively, exceeding those of several more complex methods. The learned embeddings are also robust across image classes, as further confirmed by the qualitative and class-wise analyses. Overall, the results confirm that compact metric-learning representations provide a practical, scalable, and efficient approach for contemporary CBIR applications.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hersh M. Hama, Kamaran H. Manguri, Saman M. Omer (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

















