并利用该模型推断目标文本的潜在语义结构信息; ② 通过定义语义段落内凝聚性和语义段落间发散性2个目标函数
将文本分割问题转化为多目标优化问题.采用一种针对文本分割的并行遗传算法
获得全局最优解.通过实验
在文本数据稀疏的情况下
该算法在准确率方面优于多元判别分析(MDA)方法和基于LDA的文本分割方法
对于提高文本分割的准确率是可行和有效的.
Abstract
Focusing on the data sparseness of short texts
an algorithm based on knowledge from external corpus is proposed to improve the accuracy of text segmentation
which contains two steps: Gibbs sampling is adopted to estimate the LDA model; corresponding to the corpus and the latent semantic structure information of the text is inferred based on the LDA model. Two objective functions of internal cohesion and external dissimilarity are then defined to transform text segmentation into a multi-objective optimization problem. A parallel genetic algorithm based on the objective functions is emloyed to obtain the global optimal solution for text segmentation. According to the experiments
the proposed algorithm achieves higher accuracy than the MDA and LDA-based methods in the case of data sparseness.
SHI Jing, HU Ming, SHI Xin, et al. Text segmentation based on model LDA [J]. Chinese Journal of Computers, 2008, 31(10): 1865-1873.
BRANTS T, CHEN F, TSOCHANTARIDIS I. Topic-based document segmentation with probabilistic latent semantic analysis [C]∥Proceedings of the 11th International Conference on Information and Knowledge Management. McLean, Virginia, USA: ACM Press, 2002: 211-218.
SHI Jing, DAI Guozhong. Text segmentation based on PLSA model [J]. Journal of Computer Research and Development, 2007, 44(2): 242-248.
BLEI D M, ANDREW Y N, JORDAN M I. Latent Dirichlet allocation [J]. Journal of Machine Learning Research, 2003(3): 993-1022.
GRIFFITHS T, STEYVERS M. Finding scientific topics [C]∥Proceedings of the National Academy of Sciences.Washington, USA:National Academy of Sciences,2004: 5228-5235.
ZITZLER E. Evolutionary algorithms for multiobjective optimization: methods and applications [D]. Zurich, Swiss: Swiss Federal Institute of Technology Zurich, 1999.
WITTEN I H, FRANK E. Data mining: practical machine learning tools and techniques [M]. Singapore: Elsevier, 2005.
JUNG Y. Design and evaluation of clustering criterion for optimal hierarchical agglomerative clustering [D]. Twin Cities,USA: University of Minnesota, 2001.
PEVZNER L, HEARST M. A critique and improvement of an evaluation metric for text segmentation [J]. Computational Linguistics, 2002(28): 1-19.
BEEFERMAN D, BERGER A, LAFFERTY J. Text segmentation using exponential models [C]∥Proceedings of the 2nd Conference on Empirical Methods in Natural Language Processing. Rhode Island, USA: EMNLP, 1997: 35-46.