An image retrieval method based on deep convolutional features of joint weighting aggregation is proposed to solve the problem that most existing image retrieval approaches can't fully extract image features and their performance requires to be improved. Firstly
the method extracts the outputs of the last convolutional layer as deep convolutional features of an image by passing the image through a pre-trained deep convolutional neural network. Then
the spatial weight matrix is calculated to highlight significance regions of the image and to suppress the background noise of the image. The maximum-principle of channel variance is then used to select the corresponding feature map and to calculate the spatial weight matrix. The original deep convolutional features are weighted and aggregated into a feature vector. Moreover
the channel weight vector is calculated by distinguishing feature maps of different channels
and then the global feature representation of this image is obtained by multiplying the aggregated feature vector and the channel weight. Experimental results on different public available datasets for image retrieval show that the proposed approach effectively enhances the discriminative ability of image features
outperforms the state-of-the-art approaches based on pre-trained networks and can be effectively applied to related fields of image retrieval.
XU Siyu, CAI Jiani, ZHU Jihua, et al. An adaptive hashing retrieval method of images based on multi-bit quantization [J]. Journal of Xi'an Jiaotong University, 2017, 51(8): 19-25.
LOWE D G. Distinctive image features from scale-invariant keypoints [J]. International Journal of Computer Vision, 2004, 60(2): 91-110.
JEGOU H, PERRONNIN F, DOUZE M, et al. Aggregating local image descriptors into compact codes [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2012, 34(9): 1704-1716.
JEGOU H, ZISSERMAN A. Triangulation embedding and democratic aggregation for image search [C]∥Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2014: 3310-3317.
JEGOU H, DOUZE M, SCHMID C, et al. Aggregating local descriptors into a compact image representation [C]∥Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2010: 3304-3311.
PERRONNIN F, LIU Y, SANCHEZ J, et al. Large-scale image retrieval with compressed Fisher vectors [J]. Computer Vision and Pattern Recognition, 2010, 26(2): 3384-3391.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks [C]∥Proceedings of the International Conference on Neural Information Processing Systems. Cambridge, MA, USA: MIT Press, 2012: 1097-1105.
WEI Xiushen, LUO Jianhao, WU Jianxin, et al. Selective convolutional descriptor aggregation for fine-grained image retrieval [J]. IEEE Transactions on Image Processing, 2016, 26(6): 2868-2881.
RAZAVIAN A S, SULLIVAN J, CARLSSON S, et al. Visual instance retrieval with deep convolutional networks [J]. ITE Transactions on Media Technology and Applications, 2016, 4(3): 251-258.
ZHENG Liang, YANG Yi, TIAN Qi. SIFT meets CNN: a decade survey of instance retrieval [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(5): 1224-1244.
BABENKO A, SLESAREV A, CHIGORIN A, et al. Neural codes for image retrieval [C]∥Proceedings of the 13th European Conference on Computer Vision. Berlin, Germany: Springer, 2014: 584-599.
YANDEX A B, LEMPITSKY V. Aggregating local deep features for image retrieval [C]∥Proceedings of the 2016 IEEE International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2016: 1269-1277.
KALANTIDIS Y, MELLINA C, OSINDERO S. Cross-dimensional weighting for aggregated deep convolutional features [C]∥Proceedings of the 14th European Conference on Computer Vision. Berlin, Germany: Springer-Verlag, 2016: 685-701.
TOLIAS G, SICRE R, JÉGOU H. Particular object retrieval with integral max-pooling of CNN activations[EB/OL].(2016-02-24)[2017-08-20]. https: ∥arxiv. org/abs/1511.05879v2.
WANG Jiaxing, ZHU Jihua, PANG Shanming, et al. Adaptive co-weighting deep convolutional features for object retrieval[EB/OL].(2018-03-20)[2018-04-20]. https: ∥arxiv. org/abs/1803.07360.
JEGOU H, DOUZE M, SCHMID C. On the burstiness of visual elements [C]∥Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2009: 1169-1176.
XU Jian, WANG Chunheng, QI Chengzuo, et al. Unsupervised semantic-based aggregation of deep convolutional features [J]. IEEE Transactions on Image Processing, 2018, 28(2): 601-611.
SIMONYAN K, ZISSERMAN A. Very deep convolutional networks for large-scale image recognition[EB/OL].(2015-04-10)[2017-07-20]. https: ∥arxiv. org/abs/1409.1556.
PHILBIN J, CHUM O, ISARD M, et al. Object retrieval with large vocabularies and fast spatial matching [C]∥Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2007: 1-8.
JAMES P, ONDREJ C, MICHAEL I, et al. Lost in quantization: improving particular object retrieval in large scale image databases [C]∥Proceedings of the 26th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2008: 1-8.
JEGOU H, DOUZE M, SCHMID C. Improving bag-of-features for large scale image search [J]. International Journal of Computer Vision, 2010, 87(3): 316-336.
ARANDJELOVIC R, GRONAT P, TORII A, et al. NetVLAD: CNN architecture for weakly supervised place recognition [C]∥Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2016: 5297-5307.