A Continuous Speech Recognition Method Using Dependent Feature Transformation and Combination of Subspace Region[J]. 2016, 50(4): 60-67.
DOI:
A Continuous Speech Recognition Method Using Dependent Feature Transformation and Combination of Subspace Region[J]. 2016, 50(4): 60-67.DOI: 10.7652/xjtuxb201604010.
A Continuous Speech Recognition Method Using Dependent Feature Transformation and Combination of Subspace Region
A speech recognition method based on dependent feature transformation and combination of subspace regions(MFCC-BN-TC)is proposed to improve the recognition accuracy. The structure feature(BN)and envelope feature(MFCC)are extracted to separately describe the structure and envelope information of the short speech spectrum
and the region dependent feature transformation is adopted to perform feature transformation for the BN and the MFCC
respectively.The transformation is then generalized to give a subspace region-dependent feature transformation so that different time units(frame and segment)are applied to finish multi-level modeling. Moreover
a feature combination framework is proposed
and the acoustic model is trained using combined multi-features after transformation. Experimental results and comparisons with the method using raw BN and the method based on MFCC feature show that the recognition rate of the MFCC-BN-TC method increases by 0.96% and 1.62%
respectively. The gain in performance of the MFCC-BN-TC method increases by 1.5% through combining the transformed features.
关键词
Keywords
references
NASERSHARIF B, AKBARI A. SNR-dependent compression of enhanced Mel subband energies for compensation of noise effects on MFCC features [J]. Pattern Recognition Letters, 2011, 28(11): 1320-1326.
LIU Xiaoming, BAN Chaofan, FENG Xiaorong. A short time spectrum estimation algorithm of speech enhancement under the distortion control [J]. Journal of Xi'an Jiaotong University, 2011, 45(8): 78-84.
POVEY D, KINGSBURY B, MANGU L, et al. fMPE: Discriminatively trained features for speech recognition [C]∥Proceedings of the International Conference on Audio, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2005: 961-964.
ZHANG B, MATSOUKAS S, SCHWARTZ R. Recent progress on the discriminative region-dependent transform for speech feature extraction [C]∥Proceedings of the Annual Conference of International Speech Communication Association. Baixs, France: ISCA, 2006: 1495-1498.
YAN Z, HUO Q, XU J, et al. Tied-state based discriminative training of context-expanded region-dependent feature transforms for LVCSR [C]∥Proceedings of the International Conference on Audio, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2013: 6940-6944.
YUAN Shenglong, GUO Wu, DAI Lirong. Speech recognition based on deep neural networks on Tibetan corpus [J]. Pattern Recognition and Artificial Intelligence, 2015, 28(3): 209-213.
SAINATH T N, KINGSBURY B, RAMABHADRAN B. Auto-encoder bottleneck features using deep belief networks [C]∥Proceedings of the International Conference on Audio, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2012: 4153-4156.
SAON G, KINGSBURY B. Discriminative feature-space transforms using deep neural networks [C]∥Proceedings of the Annual Conference of International Speech Communication Association. Baixs, France: ISCA, 2012: 14-17.
PAULIK M. Lattice-based training of bottleneck feature extraction neural networks [C]∥Proceedings of the Annual Conference of International Speech Communication Association. Baixs, France: ISCA, 2013: 89-93.
LIU D Y, WEI S, GUO W, et al. Lattice based optimization of bottleneck feature extractor with linear transformation [C]∥Proceedings of the International Conference on Audio, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2014: 5617-5621.
YU D, SELTZER M L. Improved bottleneck features using pretrained deep neural networks [C]∥Proceedings of the Annual Conference of International Speech Communication Association. Baixs, France: ISCA, 2011: 237-240.
HOFFMEISTER B, KLEIN T, SCHLUTER R, et al. Frame based system combination and a comparison with weighted ROVER and CNC [C]∥Proceedings of the International Conference on Spoken Language Processing. Piscataway, NJ, USA: IEEE, 2006: 537-540.
ZOU H, HASTIE T. Regularization and variable selection via the elastic net [J]. Journal of the Royal Statistical Society: Series B Statistical Methodology, 2005, 67(2): 301-320.
BECK A, TEBOULLE M. A fast iterative shrinkage thresholding algorithm for linear inverse problems [J]. SIAM Journal on Imaging Sciences, 2009, 2(1): 183-202.[16] CHEN X, ZHAO Y. Building acoustic model ensembles by data sampling with enhanced trainings and features [J]. IEEE Transactions on Audio Speech and Language Processing, 2013, 21(3): 498-507.
POVEY D, KANEVSKY D, KINGSBURY B, et al. Boosted MMI for model and feature space discriminative training [C]∥Proceedings of the International Conference on Acoustics, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2008: 4057-4060.