昆明理工大学信息工程与自动化学院,昆明,650051
网络首发:2011-10-10,
纸质出版:2011
移动端阅览
张涛 1, 余正涛 1, 郭剑毅 1, 等. 融合特征约束模型的纳西汉语双语词语对齐算法[J]. 西安交通大学学报, 2011,45(10):48-53.
A Bilingual Word Alignment Algorithm of Naxi-Chinese Based on Feature Constraint Models[J]. 2011, 45(10): 48-53.
针对纳西语、汉语因句法结构差异较大而导致双语词语自动对齐较为困难的问题
提出一种融合特征约束模型的纳西-汉语双语词语对齐算法.首先在语料中统计纳西-汉语词语区间扭曲和位置转换特性
并由此建立2个双语词语对齐的特征约束模型; 然后将提出的特征约束模型融入词语对齐的对数线性模型框架
并结合最小错误率算法训练模型参数; 最终搜索出最佳的词语对齐结果.实验以IBM Model3为词语对齐比较模型
结果表明
该双语词语对齐算法可以使纳西-汉语词语的对齐准确率提升21.9%.
A bilingual word alignment algorithm of Naxi-Chinese based on feature constraint models is proposed to reduce the difficulty of bilingual word alignment for Naxi-Chinese which has huge difference in syntactic structure. Two feature constraint models- interval distortion model and position transformation model are established by counting the traits of interval distortion and position transformation in corpus
and are integrated into a log-linear framework of word alignment. Then parameters in the models are trained using the minimum error rate algorithm and the best alignment results are eventually searched. Experimental results on IBM Model3 show that the proposed algorithm increases the word alignment accuracy of Naxi-Chinese about 21.9%.
BROWN P F, PIETRA D V J, PIETRA D S A, et al. The mathematics of statistical machine translation: parameter estimation [J].Computational Linguistics,1993, 19(2):263-311.
VOGEL S, NEY H, TILLMANN C. HMM-based word alignment in statistical translation [C]∥Proceedings of the 16th International Conference on Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 1996: 836-841.
TASKAR B, LACOSTE-JULIEN S, KLEIN D. A discriminative matching approach to word alignment [C]∥Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Porcessing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2005: 73-80.
MOORE R. A discriminative framework for bilingual word alignment [C]∥Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2005: 81-88.
CHERRY C,LIN D. A probability model to improve word alignment [C]∥Proceedings of the 41st Annual Meeting on Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2003: 88-95.
LIU Yang, LIU Qun, LIN Shouxun. Log-linear models for word alignment[C]∥Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2005: 459-466.
LIU Yang, LIU Qun, LIN Shouxun. Discriminative word alignment by linear modeling [J]. Computational Linguistics, 2010, 36(3): 303-339.
AYAN N F, DORR B J. A maximum entropy approach to combining word alignments [C]∥Proceedings of the Human Language Technology Conference of the North American Chapter of the ACL. Stroudsburg, PA, USA: Association for Computational Linguistics, 2006: 96-103.
OCH F J, NAY H. Discriminative training and maximum entropy models for statistical machine translation [C]∥Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2002: 295-302.
刘洋.树到串统计翻译模型研究[D].北京:中国科学院计算技术研究所,2007.
TOUTANOVA K, TOLAG I H, MANNING C D. Extensions to HMM-based statistical word alignment models [C]∥Proceedings of the Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2002: 87-94.
OCH F J. Minimum error rate training in statistical machine translation [C]∥Proceedings of the 41st Annual Meeting of the ACL. Stroudsburg, PA, USA: Association for Computational Linguistics, 2003:160-167.
0
浏览量
4
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621