A self-learning method of new pronunciation lexicons based on a hybrid speech recognition system is proposed to solve the problem that the existing self-expanding methods of pronunciation lexicons can only learn new words from text data but cannot learn from audio data. The method utilizes both the syllables and the graphones hybrid systems to recognize the out-of-vocabulary words in the audio data and then obtains as many new words with their pronunciations as possible by using the complementary information of the two systems. Then the new word and its pronunciation candidates are optimized using a perceptron model and a maximum entropy model to reduce the error rate. Finally
the lexicon is expanded and the language model parameters are updated by using syntactic and semantic information. Experimental results of continuous speech recognition on Wall Street Journal speech database show that the proposed method learns new words from audio data effectively
and the accuracy is greatly improved by using the data optimization strategies. The extended lexicon system yields a relative gain of 13.4% over the base line system in terms of word error rates.
关键词
Keywords
references
DAVEL M, MARTIROSIAN O. Pronunciation diction-nary development in resource-scarce environments [C]∥Proceedings of International Speech Communication Association. Grenoble, France: ISCA, 2009: 2851-2854.
BISANI M, NEY H. Joint-sequence models for grapheme-to-phoneme conversion [J]. Speech Communication, 2008, 50(5): 434-451.
RAO K, PENG F, SAK H, et al. Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks [C]∥Proceedings of International Conference on Acoustics, Speech, and Signal Processing. Piscataway, NJ, USA: IEEE, 2015: 4225-4229.
TIM S, OCHS S, TANJA S. Web-based tools and methods for rapid pronunciation dictionary creation [J]. Speech Communication, 2014, 56(1): 101-118.
BERT R, KRIS D, MARTENS J. An improved two-stage mixed language model approach for handling out-of-vocabulary words in large vocabulary continuous speech recognition [J]. Computer Speech and Language, 2014, 28(1): 141-162.
ZHENG Tieran, HAN Jiqing, LI Haiyang. Study on performance optimization for Chinese speech retrieval [J]. Journal on Communications, 2009, 30(3): 84-88.
HE Y Z, BRIAN H, PRTER B. Subword-based modeling for handling OOV words in keyword spotting [C]∥Proceedings of International Conference on Acoustics, Speech, and Signal Processing. Piscataway, NJ, USA: IEEE, 2014: 7914-7918.
QIN L, RUDNICKY A I. OOV word detection using hybrid models with mixed types of fragments [C]∥Proceedings of International Speech Communication Association. Grenoble, France: ISCA, 2012: 2450-2453.
BASHA S, AMR M, HAHN S. Improved strategies for a zero OOV rate LVCSR system [C]∥Proceedings of International Conference on Acoustics, Speech, and Signal Processing. Piscataway, NJ, USA: IEEE, 2015: 5048-5052.
BLACK A W, TAYLOR P, CALEY R. The festival speech synthesis system [EB/OL].(2002-12-27)[2016-01-04]. http: ∥www.festvox.org/docs/manual-1.4.3/.
HAN Bing, LIU Yijia, CHE Wanxiang. An incremental learning scheme for perceptron based Chinese segmentation [J]. Journal of Chinese Information, 2015, 29(5): 49-54.
李素建, 王厚峰, 俞士汶. 关键词自动标引的最大熵模型应用研究 [J]. 计算机学报, 2004, 27(9): 1192-1197.LI Sujian, WANG Houfeng, YU Shiwen. Research on maximum entropy model for keyword indexing [J]. Chinese Journal of Computers, 2004, 27(9): 1192-1197.
KLEIN D, MANNING C. Feature-rich part-of-speech tagging with a cyclic dependency network [C]∥Proceedings of Human Language Technology and North American Chapter of the Association for Computational Linguistics. Cambridge, MA, USA: ACL, 2003: 252-259.
MILLER G. WordNet: a lexical database for English [J]. Communications of the ACM, 1995, 38(11): 39-41.
Carnegie Mellon University. The CMU pronunciation dictionary [EB/OL].(2007-03-19)[2016-01-04]. http: ∥www.speech.cs.cmu.edu/cgi-bin/cmudict.