To obtain the long-distance dependence between different feature points in the feature map of the convolutional neural network
so that the convolutional neural network can better distinguish the foreground target and background information
a spatial attention mechanism with global features is proposed. Combining the multi-channel original feature map into a single-channel feature fusion map through the channel fusion layer
the influence of the information distribution between channels on obtaining spatial attention weights is eliminated; the feature fusion map through global feature acquisition processing to obtain the global feature map
which represents the correlation between a feature point and all points in the feature fusion map; the global feature map is multiplied by the learnable variable with an initial value of 0
and copies itself to the size of the original feature map in channel domain. The expanded global feature map is added to the original feature map to obtain a feature map with attention mechanism. After adding the spatial attention mechanism with global features to different convolutional neural networks
the experimental results show that in the brain wave two-classification task
the algorithm classification accuracy is heightenedby the maximum of 0.839%; in the CIFAR-10 data set multi-classification task
the classification accuracy of the algorithm is improvedby the maximum of 0.484%; in the single-category detection of night vehicles
the algorithm is heightenedby the maximum of 3.860% with the evaluation standard of the average accuracy value of Intersection over union(IoU)greater than 0.5
and the algorithm is improvedby the maximum of 11.726% with the evaluation standard of the average accuracy value of IoU greater than 0.75; in the multi-category detection of the voc2007 data set
the algorithm has improved by the maximum of 0.778% with the evaluation criterion of the average precision value of IoU greater than 0.5
and the algorithm has improvedby the maximum of 1.232% with the evaluation criterion of the average precision value of IoU greater than 0.75.
关键词
Keywords
references
GEHRING J, AULI M, GRANGIER D, et al. A convolutional encoder model for neural machine translation [C]∥Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics: Volume 1 Long Papers. Stroudsburg, PA, USA: Association for Computational Linguistics, 2017: 123-135.
GEHRING J, AULI M, GRANGIER D, et al. Convolutional sequence to sequence learning [C]∥Proceedings of the 34th International Conference on Machine Learning. Princeton, NJ, USA: International Machine Learning Society(IMLS), 2017: 2029-2042.
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need [C]∥Proceedings of the 31st Annual Conference on Neural Information Processing Systems. Vancouver, Canada: NIPS, 2017: 5999-6009.
WANG F, JIANG M Q, QIAN C, et al. Residual attention network for image classification [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 6450-6458.
HU H, GU J Y, ZHANG Z, et al. Relation networks for object detection [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 3588-3597.
HUANG Z L, WANG X G, HUANG L C, et al. CCNet: criss-cross attention for semantic segmentation [C]∥Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2019: 603-612.
YUAN Yuhui, WANG Jingdong. OCNet: object context network for scene parsing [EB/OL].(2019-01-22)[2019-12-06].https:∥arxiv.org/abs/1809.00916.
HU J, SHEN L, SUN G. Squeeze-and-excitation networks [C]∥Proceedings of the 2018 IEEE/CVF Con-ference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7132-7141.
CAO Yue, XU Jiarui, LIN S, et al. GCNet: non-local networks meet squeeze-excitation networks and beyond [EB/OL].(2019-04-25)[2019-12-06]. https: ∥arxiv.org/abs/1904.11492.
WANG X L, GIRSHICK R, GUPTA A, et al. Non-local neural networks [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7794-7803.
WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]∥Proceedings of the 15th European Conference on Computer Vision. Cham, Germany: Springer, 2018: 3-19.
ROY A G, NAVAB N, WACHINGER C. Concurrent spatial and channel ‘squeeze excitation’ in fully convolutional networks [C]∥Proceedings of the 21st International Conference on Medical Image Computing and Computer Assisted Intervention. Cham, Germany: Springer, 2018: 421-429.
BUADES A, COLL B, MOREL J M. A non-local algorithm for image denoising [C]∥Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2005: 60-65.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60(6): 84-90.
SIMONYAN K, ZISSERMAN. Very deep convolutional networks for large-scale image recognition [EB/OL].(2015-04-10)[2019-12-06]. https: ∥arxiv.org/abs/1409.1556.
REDMON J, FARHADI A. Yolov3: an incremental improvement [EB/OL].(2018-04-08)[2019-12-06]. https: ∥arxiv.org/abs/1804.02767.
REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]∥Proceedings of the 29th Annual Conference on Neural Information Processing Systems. Vancouver, Canada: NIPS, 2015: 91-95.