1.长安大学信息工程学院,710064,西安
2.西安电子科技大学电子工程学院,710071,西安
收稿:2025-10-13,
修回:2026-03-30,
录用:2026-04-27,
移动端阅览
刘妮, 雷娇娇, 邢振蓉, 等. 数据增强的深度卷积网络不变性机理的可解释性研究[J/OL]. 西安交通大学学报, 2026.
LIU Ni, LEI Jiao Jiao, XING ZHEN RONG, et al. Interpreting the Invariance Mechanism of Data Augmentation on Deep Convolutional Networks[J/OL]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY, 2026.
在图像分类任务中,数据增强是提升深度卷积网络(ConvNets)分类不变性的通用手段。针对当前深度卷积网络不变性研究中主要聚焦网络层的整体不变性度量,而忽视了位置、方向等信息的表征机制,且缺乏对特征综合机理的系统阐释的问题,本文受认知神经科学领域相关研究的启发,采用圆形/方形二分类任务构建简化实验场景,有效规避复杂特征的干扰。采用最大激活方法,对全连接层与全局平均池化(AvgPool)层进行特征可视化分析,研究了经数据增强训练后的全连接层与AvgPool层的不变性学习特性,并采用t分布随机邻域嵌入(t-SNE)降维可视化揭示了层级间高维权重向量的分类决策机理。结果表明:(1)全连接层首层通过逐步扩大位置感知域的方式实现了平移不变性;(2)全连接层通过线性加权机制有效整合分布式局部特征,遵循同类别特性正向激发、异类别特征负向抑制的连接逻辑;(3)ResNet架构中AvgPool层展现更优的平移不变性,而全连接结构则表现出更强的旋转和缩放不变性。
In image classification tasks
data augmentation is a common means to improve the classification invariance of deep convolutional networks (ConvNets ).To address the problems that current research on Convolutional Neural Network invariance primarily focuses on global metrics while neglecting the representation mechanisms of position and direction information
and lacks a systematic explanation of feature integration mechanisms
this study
inspired by relevant research in cognitive science and neuroscience
constructed a simplified experimental scenario using a circular/square binary classification task to effectively eliminate interference from complex features. By employing the maximum activation method
we visualized and analyzed the feature representation and invariance learning characteristics of the fully connected (FC) layers and Average Pooling (AvgPool) layers in models trained with data augmentation. Additionally
t-SNE dimensionality reduction visualization was used to reveal the classification decision-making mechanisms of high-dimensional weight vectors between layers. The results demonstrate that: (1) the initial fully connected layer achieves translation invariance by gradually expanding the position-aware domain; (2) the fully connected layer effectively integrates distributed local features through a linear weighting mechanism
following a connectivity logic of positive excitation for intra-class features and negative inhibition for inter-class features; and (3) the AvgPool layer in the ResNet architecture exhibits superior translation invariance
while the cascaded fully connected structure demonstrates stronger rotation and scaling invariance.
BLYTHING R , BISCIONE V , VANKOV I I , et al . The human visual system and CNNs can both support robust online translation tolerance following extreme displacements [J ] . Journal of Vision , 2021 , 21 ( 2 ): 9 - 9 .
LECUN Y , BOTTOU L , BENGIO Y , et al . Gradient-based learning applied to document recognition [J ] . Proceedings of the IEEE , 2002 , 86 ( 11 ): 2278 - 2324 .
JADERBERG M , SIMONYAN K , ZISSERMAN A . Spatial transformer networks [C ] // Advances in Neural Information Processing Systems . 2015 .
BISCIONE V , BOWERS J . Learning translation invariance in CNNs [J ] . arXiv preprint arXiv: 2011.11757 , 2020 .
LAPTEV D , SAVINOV N , BUHMANN J M , et al . Ti-pooling: Transformation-invariant pooling for feature learning in convolutional neural networks [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2016 : 289 - 297 .
顾正强 , 张严 , 张冰尘 . 一种基于模糊滤波提高SAR自动目标识别平移不变性的方法 [J ] . 系统工程与电子技术 , 2020 .
GU ZhengQiang , ZHANG Yan , ZHANG Bingchen . Improving translation invariance of SAR automatic target recognition based on blur filtering method [J ] . Systems Engineering and Electronics , 2020 , 42 ( 11 ): 2488 - 2496 .
李俊英 . 深度卷积神经网络的旋转等变性研究 [D ] . 浙江大学 , 2019 .
LI Junying . Research on Rotation Equivariance of Deep Convolutional Neural Networks [D ] . Zhejiang University , 2019 .
KRIZHEVSKY A , SUTSKEVER I , HINTON G E . Imagenet classification with deep convolutional neural networks [J ] . Advances in Neural Information Processing Systems , 2012 , 25 .
HERMANN K , CHEN T , KORNBLITH S . The origins and prevalence of texture bias in convolutional neural networks [C ] // Proceedings of the Advances in Neural Information Processing Systems . 2020 .
MYBURGH J C , MOUTON C , DAVEL M H . Tracking translation invariance in CNNs [C ] // Proceedings of the Southern African Conference for Artificial Intelligence Research . Springer , 2020 : 282 – 295 .
KAUDERER-ABRAMS E . Quantifying translation-invariance in convolutional neural networks [J ] . arXiv preprint arXiv:1801. 01450 , 2017
POSPISIL D A , PASUPATHY A , BAIR W . ‘Artiphysiology’ reveals V4-like shape tuning in a deep network trained for image classification [J ] . Elife , 2018 , 7 : e38242 .
BAKER N , LU H , ERLIKHMAN G , et al . Local features and global shape information in object classification by deep convolutional neural networks [J ] . Vision Research , 2020 , 172 : 46 - 61 .
ERHAN D , BENGIO Y , COURVILLE A , et al . Visualizing higher-layer features of a deep network [J ] . University of Montreal , 2009 , 1341 ( 3 ): 1 .
YOSINSKI J , CLUNE J , NGUYEN A , ET al . Understanding neural networks through deep visualization [J ] . arXiv preprint arXiv: 1506.06579 , 2015 .
NGUYEN A , YOSINSKI J , CLUNE J . Understanding neural networks via feature visualization: A survey [M ] // Explainable AI: Interpreting, Explaining and Visualizing Deep Learning . Springer , 2019 : 55 - 76 .
RIBEIRO M T , SINGH S , GUESTRIN C . “ Why should i trust you?” explaining the predictions of any classifier [C ] // Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 2016 : 1135 - 1144 .
MAATEN L VAN DER , HINTON G . Visualizing data using t-SNE [J ] . Journal of Machine Learning Research , 2008 , 9(Nov): 2579- 2605 .
DENG L . The MNIST database of handwritten digit images for machine learning research [best of the web] [J ] . IEEE Signal Processing Magazine , 2012 , 29 ( 6 ): 141 - 142 .
KAYHAN O S , GEMERT J C VAN . On translation invariance in cnns: Convolutional layers can exploit absolute spatial location [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2020 : 14274 - 14285 .
HAN Y , ROIG G , GEIGER G , et al . Scale and translation-invariance for novel objects in human vision [J ] . Scientific Reports , 2020 , 10 ( 1 ): 1411 .
ZEILER M D , FERGUS R . Visualizing and understanding convolutional networks [C ] // Proceedings of the European Conference on Computer Vision . Springer , 2014 : 818 – 833 .
ZHOU B , KHOSLA A , LAPEDRIZA A , et al . Learning deep features for discriminative localization [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2016 : 2921 - 2929 .
尚骏远 , 杨乐涵 , 何琨 . 基于特征可视化分析深度神经网络的内部表征 [J ] . 计算机科学 , 2020 , 47 ( 5 ): 190 - 197 .
SHANG Junyuan , YANG Lehan , HE Kun . Analyzing Internal Representations of Deep Neural Networks Based on Feature Visualization [J ] . Computer Science , 2020 , 47 ( 5 ): 190 - 197 .
ZHANG Q , YANG Y , MA H , et al . Interpreting cnns via decision trees [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2019 : 6261 - 6270 .
HU J , GAO J , YE J , et al . Model LEGO: Creating models like disassembling and assembling building blocks [C ] // Proceedings of the Advances in Neural Information Processing Systems . 2024 .
NGUYEN A , YOSINSKI J , CLUNE J . Deep neural networks are easily fooled: High confidence predictions for unrecognizable images [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2015 : 427 - 436 .
SZEGEDY C , ZAREMBA W , SUTSKEVER I , et al . Intriguing properties of neural networks [J ] . arXiv preprint arXiv: 1312.6199 , 2013 .
WANG F , LIU H , CHENG J . Visualizing deep neural networks by alternately image blurring and deblurring [J ] . Neural Networks , 2018 , 97 : 162 - 172 .
王芳珍 , 张小丽 , 赵琦武 , 等 . 一维卷积神经网络在机械故障特征提取中的可解释性研究 [J ] . 西安交通大学学报 , 2025 , 59 ( 7 ): 24 - 35 .
WANG Fangzhen , ZHAGN Xiaoli , ZHAO Qiwu , et al . Study on the Interpretability of One-Dimensional Convolutional Neural Networks in Mechanical Fault Feature Extraction [J ] . Journal of Xi’an Jiaotong University , 2025 , 59 ( 7 ): 24 - 35 .
ROUSSEEUW P J . Silhouettes: A graphical aid to the interpretation and validation of cluster analysis [J ] . Journal of Computational and Applied Mathematics , 1987 , 20 : 53 - 65 .
JANSSON Y , LINDEBERG T . Scale-invariant scale-channel networks: Deep networks that generalise to previously unseen scales [J ] . Journal of Mathematical Imaging and Vision , 2022 , 64 ( 5 ): 506 - 536 .
0
浏览量
1
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621