To solve the problem of gradient dispersion in stacking auto-encoder(SAE)with large number of parameters
we analyze the distribution of encoding values in each hidden-layer of network. It is found that most of them are distributed in the saturation area of the activation function
which directly leads to a weight gradient loss
thus a normalizing strategy is introduced. The node is normalized according to the sample
then two parameters are introduced to scale and move the encoding values. Then the modified values are passed to the activation function to next layer. The fault diagnosis of rolling bearing is carried out by the normalized SAE
and the spectrum of vibration signal is input into the network. Compared with ordinary SAE
the encoding values of normalized SAE are more well-distributed. For example
the entropy of encoding values in the first level is increased from 0.88 bit to 16.29 bit. The normalized SAE has higher anti-noise ability and faster training rate. When the signal to noise ratio(SNR)is 0 dB
the recognition accuracy is increased from 16.18% to 100% on the rolling bearing data sets of Case Western Reserve University. On laboratory data sets
the training time is decreased by 37.22%
and the recognition accuracy is increased from 97.93% to 99.95%. The introduced normalizing strategy provides a reference for subsequent research on the construction of SAE
and also provides a strategy for fault diagnosis of rolling bearings.
LEI Yaguo, JIA Feng, KONG Detong, et al. Opportunities and challenges of machinery intelligent fault diagnosis in big data era [J]. Chinese Journal of Mechanical Engineering, 2018, 54(5): 94-104.
REN Hao, QU Jianfeng, CHAI Yi, et al. Deep learning for fault diagnosis: the state of the art and challenge [J]. Control and Decision, 2017(8): 1345-1358.
LEI Yaguo, JIA Feng, ZHOU Xin, et al. A deep learning-based method for machinery health monitoring with big data [J]. Chinese Journal of Mechanical Engineering, 2015, 51(21): 49-56.
ZHANG Qingchen, YANG L T, CHEN Zhikui. Deep computation model for unsupervised feature learning on big data [J]. IEEE Transactions on Services Computing, 2016, 9(1): 161-171.
RUMELHART D E, HINTON G E, WILLIAMS R J. Learning representations by back-propagating errors [J]. Nature, 1988, 323(6088): 399-421.
HINTON G E, SALAKHUTDINOV R R. Reducing the dimensionality of data with neural networks [J]. Science, 2006, 313(5786): 504-507.
SCHÖLKOPF B, PLATT J, HOFMANN T. Greedy layer-wise training of deep networks [C]∥International Conference on Neural Information Processing Systems. Cambridge, MA, USA: MIT Press, 2006: 153-160.
VINCENT P, LAROCHELLE H, BENGIO Y, et al. Extracting and composing robust features with denoising autoencoders [Z]. New York, USA: ACM, 2008: 1096-1103.
AMARAL T, KANDASWAMY C, SANTOS J M. Using different cost functions to train stacked auto-encoders [C]∥Mexican International Conference on Artificial Intelligence. Piscataway, NJ, USA: IEEE, 2014: 114-120.
JIA Feng, LEI Yaguo, LIN Jing, et al. Deep neural networks: a promising tool for fault characteristic mining and intelligent diagnosis of rotating machinery with massive data [J]. Mechanical Systems and Signal Processing, 2016, 72/73(1): 303-315.
SHAO H, JIANG H, ZHAO H, et al. A novel deep autoencoder feature learning method for rotating machinery fault diagnosis [J]. Mechanical Systems and Signal Processing, 2017, 95(3): 187-204.
CUI Jiang, TANG Junxiang, GONG Chunying, et al. A fault feature extraction method of aerospace generator rotating rectifier based on improved stacked auto-encoder [J]. Proceedings of the CSEE, 2017, 37(19): 5696-5706.
WANG L, ZHAO X, PEI J, et al. Transformer fault diagnosis using continuous sparse autoencoder [J]. SpringerPlus, 2016, 5(1): 448-463.
SHAHERYAR A, YIN X, HAO H, et al. A denoising based autoassociative model for robust sensor monitoring in nuclear power plants [J]. Science and Technology of Nuclear Installations, 2016, 3(1): 1-17.
YANG Zhixin, WANG Xianbo, ZHONG Jianhua. Representational learning for fault diagnosis of wind turbine equipment: a multi-layered extreme learning machines approach [J]. Energies, 2016, 9: 379.
IOFFE S, SZEGEDY C. Batch normalization: accelerating deep network training by reducing internal covariate shift [C]∥32nd International Conference on Machine Learning. Cambridge, MA, USA: IMLS, 2015: 448-456.
HINTON G E, SALAKHUTDINOV R R. Reducing the dimensionality of data with neural networks [J]. Science, 2006, 313(5786): 504-507.
KINGMA D P, BA J L. Adam: a method for stochastic optimization [EB/OL].[2018-04-01]. http: ∥arxiv.org/pdf/1412.6980v8.pdf.
HINTON G E, SRIVASTAVA N, KRIZHEVSKY A, et al. Improving neural networks by preventing co-adaptation of feature detectors [J]. Computer Science, 2012, 3(4): 212-223.
Case Western Reserve University Bearing Data Center. Bearing data file [EB/OL].[218-01-11]. http: ∥csegroups.case.edu/bearingdatacenter/home.