长安大学信息工程学院,710064,西安
西安电子科技大学电子工程学院,710071,西安
作者简介:刘妮(1985—),女,讲师;
何立火(通信作者),男,教授,博士生导师。
收稿:2025-10-13,
网络首发:2026-04-29,
纸质出版:2026-09-10
移动端阅览
刘妮, 雷娇娇, 邢振榕, 等. 数据增强下深度卷积网络不变性机理的可解释性研究[J]. 西安交通大学学报,2026,60 (9):218-228. https://doi.org/10.7652/xjtuxb202609021.
LIU Ni, LEI Jiaojiao, XING Zhenrong, et al. Interpretability Study of the Invariance Mechanism of Deep Convolutional Networks under Data Augmentation[J]. Journal of Xi'an Jiaotong University,2026,60 (9):218-228. https://doi.org/10.7652/xjtuxb202609021.
刘妮, 雷娇娇, 邢振榕, 等. 数据增强下深度卷积网络不变性机理的可解释性研究[J]. 西安交通大学学报,2026,60 (9):218-228. https://doi.org/10.7652/xjtuxb202609021. DOI:
LIU Ni, LEI Jiaojiao, XING Zhenrong, et al. Interpretability Study of the Invariance Mechanism of Deep Convolutional Networks under Data Augmentation[J]. Journal of Xi'an Jiaotong University,2026,60 (9):218-228. https://doi.org/10.7652/xjtuxb202609021. DOI:
针对目前数据增强提升深度卷积网络不变性的研究缺乏对其实现不变性的机理进行阐释的问题,对其位置、方向等信息的表征机制和特征综合机理展开研究。首先,采用圆形和方形二分类任务构建简化实验场景,使用其增强数据训练网络;然后,通过最大激活方法对全连接层和全局平均池化层进行特征可视化,并采用
t
-SNE降维揭示了全连接层间高维权重向量的分类决策机理。结果表明:全连接层首层通过逐步扩大位置感知域的方式实现平移不变性;在圆方数据集平移增强实验中,Fc2层平移不变性得分为0.99,较Conv5层提升了45%;全连接层通过线性加权机制综合圆形和方形的局部特征,遵循同类别特性正向激发、异类别特征负向抑制的连接逻辑;ResNet架构中平均池化层展现更优的平移不变性,在圆方数据集增强实验中,平均池化层得分为0.93,高于全连接层的0.79,而级联全连接结构则表现出更强的旋转和尺度不变性。以上结论在MNIST数据集的扩展实验上也得到了验证。
Addressing the issue that current research on enhancing the invariance of deep convolutional networks through data augmentation lacks an explanation of the mechanisms underlying its implementation
this paper explores the representation and pooling mechanisms of position
orientation
and other information.First
a simplified experimental scenario is constructed using a binary classification task of circles and squares
which is employed to augment data for network training.Then
feature visualization is conducted for the fully connected layer and global average pooling layer using the maximum activation method
and
t
-SNE dimensionality reduction is adopted to reveal the classification decision-making mechanism of high-dimensional weight vectors within the fully connected layers.It is shown that translation invariance is achieved by the first fully connected layer through the gradual expansion of the position-aware receptive field.In the translation augmentation experiment on the circle-and-square dataset
a translation invariance score of 0.99 is obtained by the Fc2 layer
representing a 45% improvement over that of the Conv5 layer.The local
features of circles and squares are synthesized by the fully connected layers through a linear weighting mechanism
following a connection logic in which same-category features are positively activated and different-category features are negatively suppressed.Better translation invariance is exhibited by the average pooling layer in the ResNet architecture.In the augmentation experiment on the circle-and-square dataset
a score of 0.93 is obtained by the AvgPool layer
higher than the 0.79 obtained by the fully connected layer
whereas stronger rotation and scale invariance is exhibited by the cascaded fully connected structure.These conclusions are also verified through extended experiments on the MNIST dataset.
Blything R,Biscione V,Vankov I I,et al.The human visual system and CNNs can both support robust online translation tolerance following extreme displacements[J].Journal of Vision,2021,21(2):9.
Jaderberg M,Simonyan K,Zisserman A,et al.Spatial transformer networks[C]//Proceedings of the 28th International Conference on Neural Information Processing Systems.Cambridge,MA,USA:MIT Press,2015:2017-2025.
Biscione V,Bowers J.Learning translation invariance in CNNs[PP/OL].arXiv(2020-11-06)[2025-10-13].https://arxiv.org/abs/2011.11757.
Laptev D,Savinov N,Buhmann J M,et al.TI-POOLING:transformation-invariant pooling for feature learning in convolutional neural networks[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2016:289-297.
顾正强,张严,张冰尘.一种基于模糊滤波提高SAR自动目标识别平移不变性的方法[J].系统工程与电子技术,2020,42(11):2488-2496.
Gu Zhengqiang,Zhang Yan,Zhang Bingchen.Improving translation invariance of SAR automatic target recognition based on blur filtering method[J].Systems Engineering and Electronics,2020,42(11):2488-2496.
李俊英.深度卷积神经网络的旋转等变性研究[D].杭州:浙江大学,2019.
Krizhevsky A,Sutskever I,Hinton G E.ImageNet classification with deep convolutional neural networks[C]//Proceedings of the 26th International Conference on Neural Information Processing Systems.Red Hook, NY,USA:Curran Associates Inc.,2012:1097-1105.
Hermann K L,Chen Ting,Kornblith S.The origins and prevalence of texture bias in convolutional neural networks[C]//Proceedings of the 34th International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2020:19000-19015.
Myburgh J C, Mouton C, Davel M H.Tracking translation invariance in CNNs[C]//Artificial Intelligence Research.Cham,Switzerland:Springer International Publishing,2020:282-295.
Kauderer-Abrams E.Quantifying translation-invariance in convolutional neural networks[PP/OL].arXiv(2017-12-10)[2025-10-13].https://arxiv.org/abs/1801.01450.
Baker N,Lu Hongjing,Erlikhman G,et al.Local features and global shape information in object classification by deep convolutional neural networks[J].Vision Research,2020,172:46-61.
Erhan D,Bengio Y,Courville A,et al.Visualizing higher-layer features of a deep network[J].University of Montreal,2009,1341(3):1.
Yosinski J,Clune J,Nguyen A,et al.Understanding neural networks through deep visualization[PP/OL].arXiv(2015-06-22)[2025-10-13].https://arxiv.org/abs/1506.06579.
Nguyen A,Yosinski J,Clune J.Understanding neural networks via feature visualization:a survey[M]//Samek W,Montavon G,Vedaldi A,et al.Explainable AI:Interpreting, Explaining and Visualizing Deep Learning.Cham,Switzerland:Springer International Publishing,2019:55-76.
Ribeiro M T,Singh S,Guestrin C.“Why should I trust you?”:explaining the predictions of any classifier[C]//Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.New York,USA:ACM,2016:1135-1144.
van der Maaten L,Hinton G.Visualizing data using t -SNE[J ] .Journal of Machine Learning Research, 2008,9(86):2579-2605.
Deng Li.The MNIST database of handwritten digit images for machine learning research[J].IEEE Signal Processing Magazine,2012 ,29(6):141-142.
Semih K O,van Gemert J C.On translation invariance in CNNs:convolutional layers can exploit absolute spatial location[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2020:14262-14273.
Han Yena,Roig G,Geiger G,et al.Scale and translation-invariance for novel objects in human vision[J].Scientific Reports,2020,10(1):1411.
Zeiler M D,Fergus R.Visualizing and understanding convolutional networks[C]//Computer Vision-ECCV 2014.Cham,Switzerland:Springer International Publishing,2014:818-833.
Zhou Bolei,Khosla A,Lapedriza A,et al.Learning deep features for discriminative localization[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE, 2016:2921-2929.
Zhang Quanshi,Yang Yu,Ma Haotian,et al.Interpreting CNNs via decision trees[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2019:6254-6263.
Hu Jiacong,Gao Jing,Ye Jingwen,et al.Model LEGO:creating models like disassembling and assembling building blocks[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2024:127711-127738.
Nguyen A,Yosinski J,Clune J.Deep neural networks are easily fooled:high confidence predictions for unrecognizable images[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2015:427-436.
Szegedy C,Zaremba W,Sutskever I,et al.Intriguing properties of neural networks[PP/OL].arXiv(2014-02-19)[2025-10-13].https://arxiv.org/abs/1312.6199.
Wang Feng,Liu Haijun,Cheng Jian.Visualizing deep neural network by alternately image blurring and deblurring[J].Neural Networks,2018,97:162-172.
王芳珍,张小丽,赵琦武,等.一维卷积神经网络在机械故障特征提取中的可解释性研究[J].西安交通大学学报,2025,59(7):24-35.
Wang Fangzhen,Zhang Xiaoli,Zhao Qiwu,et al.Study on the interpretability of one-dimensional convolutional neural networks in mechanical fault feature extraction[J].Journal of Xi'an Jiaotong University, 2025,59(7):24-35.
Rousseeuw P J.Silhouettes:a graphical aid to the interpretation and validation of cluster analysis[J].Journal of Computational and Applied Mathematics, 1987,20:53-65.
Jansson Y,Lindeberg T.Scale-invariant scale-channel networks:deep networks that generalise to previously unseen scales[J].Journal of Mathematical Imaging and Vision,2022,64(5):506-536.
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621