WANG Fang, QIAO Ruiping. SPAM: Spatially Partitioned Attention Module in Deep Convolutional Neural Networks for Image Classification[J]. 2023, 57(9): 185-192.
DOI:
WANG Fang, QIAO Ruiping. SPAM: Spatially Partitioned Attention Module in Deep Convolutional Neural Networks for Image Classification[J]. 2023, 57(9): 185-192.DOI: 10.7652/xjtuxb202309019.
SPAM: Spatially Partitioned Attention Module in Deep Convolutional Neural Networks for Image Classification
Existing attention mechanisms often use fusion or compression to obtain the required information
but this leads to a large quantity of information lost in the spatial or channel dimension. In order to solve this problem
the Spatially Partitioned Attention Module(SPAM)
a really effective and lightweight attention module that can help obtain attention without channel fusion or compression
was proposed in the paper. For the input intermediate feature map
the SPAM first adaptively used average pooling and maximum pooling features for feature extraction
replaced the point feature with the local block feature in space to reduce the amount of calculation and used the IN layer and depthwise convolution to capture global spatial attention. Meanwhile
the reconstruction of channel dimension information was directly completed by group convolution. Finally
the interpolation operation was used to obtain overall attention and weight the input feature map. Notably
the SPAM can be easily embedded in various mainstream CNN architectures
and network performance can be significantly improved by increasing a few microparameters and calculations. To demonstrate the effectiveness of the SPAM
numerous experiments were conducted on the ImageNet-1K
CIFAR-100
and Food-101 datasets
and the network's regions of interest were visualized using Grad-CAM. On the ImageNet-1K
CIFAR-100
and Food-101 datasets
the SPAM improved the accuracy of the baseline network by up to about 1.08%
2.46%
and 1.09%
respectively. The results show that the performance of the network embedded with the SPAM components is greatly improved; compared to other commonly used lightweight attention mechanisms
the SPAM always works better; the SPAM can really induce the networks to pay more attention to the target object regions and accurately improve the expression ability of the networks.
关键词
Keywords
references
GU Jiuxiang, WANG Zhenhua, KUEN J, et al. Recent advances in convolutional neural networks [J]. Pattern Recognition, 2018, 77: 354-377.
LECUN Y, BOTTOU L, BENGIO Y, et al. Gradient-based learning applied to document recognition [J]. Proceedings of the IEEE, 1998, 86(11): 2278-2324.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60(6): 84-90.
LECUN Y, BENGIO Y, HINTON G. Deep learning [J]. Nature, 2015, 521(7553): 436-444.
LIU Wenxiang, SHU Yuanzhong, TANG Xiaomin, et al. Remote sensing image segmentation using dual attention mechanism Deeplabv3+ algorithm [J]. Tropical Geography, 2020, 40(2): 303-313.
LE V T, KIM Y G. Attention-based residual autoencoder for video anomaly detection [J]. Applied Intelligence, 2023, 53(3): 3240-3254.
PU Chunyu, HUANG Hong, YANG Liping. An attention-driven convolutional neural network-based multi-level spectral-spatial feature learning for hyperspectral image classification [J]. Expert Systems with Applications, 2021, 185: 115663.
HU Jie, SHEN Li, SUN Gang. Squeeze-and-excitation networks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7132-7141.
ZHANG Ke, FENG Xiaohan, GUO Yurong, et al. Overview of deep convolutional neural networks for image classification [J]. Journal of Image and Graphics, 2021, 26(10): 2305-2325.
WANG Xiaolong, GIRSHICK R, GUPTA A, et al. Non-local neural networks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7794-7803.
WOO S, PARK J, LEE J Y, et al. CBAM: convolu-tional block attention module [C]//Computer Vision-ECCV 2018. Berlin, Germany: Springer International Publishing, 2018: 3-19.
CHEN Long, ZHANG Hanwang, XIAO Jun, et al. SCA-CNN: spatial and channel-wise attention in convolutional networks for image captioning [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 6298-6306.
WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: efficient channel attention for deep convolutional neural networks [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2020: 11531-11539.
HOU Qibin, ZHOU Daquan, FENG Jiashi. Coordinate attention for efficient mobile network design [C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2021: 13708-13717.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Spatial pyramid pooling in deep convolutional networks for visual recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 37(9): 1904-1916.
HUANG Zhanchao, WANG Jianlin, FU Xuesong, et al. DC-SPP-YOLO: dense connection and spatial pyramid pooling based YOLO for object detection [J]. Information Sciences, 2020, 522: 241-258.
RUSSAKOVSKY O, DENG Jia, SU Hao, et al. ImageNet large scale visual recognition challenge [J]. International Journal of Computer Vision, 2015, 115(3): 211-252.
SELVARAJU R R, COGSWELL M, DAS A, et al. Grad-CAM: visual explanations from deep networks via gradient-based localization [C]//2017 IEEE International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2017: 618-626.
PASZKE A, GROSS S, MASSA F, et al. PyTorch: an imperative style, high-performance deep learning library [C]//Advances in Neural Information Processing Systems. San Francisco, CA, USA: Curran Associates, Inc., 2019: 8024-8035.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2016: 770-778.
RADOSAVOVIC I, KOSARAJU R P, GIRSHICK R, et al. Designing network design spaces[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2020: 10425-10433.