盲信号处理重点实验室,成都,610041
网络首发:2017-10-10,
纸质出版:2017
移动端阅览
曹卫权, 褚衍杰, 李显. 针对机器学习中残缺数据的近似补全方法[J]. 西安交通大学学报, 2017,51(10):142-148.
Approximate Imputation Method for Missing Data in Machine Learning[J]. 2017, 51(10): 142-148.
曹卫权, 褚衍杰, 李显. 针对机器学习中残缺数据的近似补全方法[J]. 西安交通大学学报, 2017,51(10):142-148. DOI: 10.7652/xjtuxb201710023.
Approximate Imputation Method for Missing Data in Machine Learning[J]. 2017, 51(10): 142-148. DOI: 10.7652/xjtuxb201710023.
针对机器学习中含残缺项的数据不能被有效利用
导致分类和回归准确率不高的问题
提出了一种近似补全方法——k-ANNO方法。给定残缺的数据样本
该方法首先通过离线构建的图结构来近似搜索与该样本最接近的k个近邻顶点
然后采用快速二次规划估计各近邻的最优权重
最后基于权重值来补全样本中的残缺项
用户可以根据实际需求在补全效率与准确性之间折中。k-ANNO方法较好地解决了机器学习中普遍存在的数据残缺问题
有效抑制了数据残缺对分类和回归精度的干扰。利用多份公开数据集评估了k-ANNO方法的补全效果
结果表明:当加速比在2~10之间时
k-ANNO方法的分类错误率比已有的均值补全、C均值补全、自组织映射补全方法低1%~4%
回归均方根误差比已有方法低约0.5~2.0; 当样本规模为4 000时
在不同加速比参数下
k-ANNO方法的计算效率比朴素k近邻方法高约35%~320%。
An approximate imputation method called k-ANNO is proposed to handle the problems of missing data in machine learning field given a missing sample. The proposed method begins by constructing an offline graph to approximately search nearest neighbors of the partially missing sample efficiently. Then a fast quadratic programming algorithm is utilized to determine the optimal weight for each neighbor. Finally
unmissed parts of the neighbors are used to impute the missing attributes by the estimated weights. Users get the freedom to weigh up between efficiency and imputation accuracy. The widespread data missing problems are well solved in this paper and k-ANNO is able to depress the impact of missing data effectively. Experiments on various well known datasets show that when the speedup rate parameters are between 2 and 10
k-ANNO method outperforms existing ones such as mean imputation or C-Means imputation etc. and the classification error and the regression error are 1% to 4% and 0.5-2.0 lower than those
respectively. Meanwhile
k-ANNO outperforms naïve k-NN imputation with a faster efficiency increased by 35%-320% faster.
杨雷, 李贵鹏, 张萍. 无线传感网络数据缺失下的通信优化仿真 [J]. 计算机仿真, 2013, 30(12): 249-252.
YANG Lei, LI Guipeng, ZHANG Ping. Simulation on communication optimization of wireless sensor networks under missing data [J]. Computer Simulation, 2013, 30(12): 249-252.
孟杰, 李春林. 基于随机森林模型的分类数据缺失值插补 [J]. 统计与信息论坛, 2014(9): 86-90.
MENG Jie, LI Chunlin. Missing data imputation for categorical data based on random forest model [J]. Statistics Information Forum, 2014(9): 86-90.
吴小姣, 李高明, 易大莉, 等. 基因表达谱的非参缺失森林填补算法研究 [J]. 中国卫生统计, 2016, 33(6): 1068-1070.
WU Xiaojiao, LI Gaoming, YI Dali, et al. Study on the algorithm of non-parameter deletion forest filling for gene expression profiles [J]. Chinese Journal of Health Statistics, 2016, 33(6): 1068-1070.
张晓琴, 程誉莹. 基于随机森林模型的成分数据缺失值填补法 [J]. 应用概率统计, 2017, 33(1): 102-110.
ZHANG Xiaoqin, CHENG Yuying. Imputation of missing values for compositional data based on random forest [J]. Chinese Journal of Applied Probability and Statistics, 2017, 33(1): 102-110.
FARHANGFAR A, KURGAN L, DY J G, et al. Impact of imputation of missing values on classification error for discrete data [J]. Pattern Recognition, 2008, 41(12): 3692-3705.
OUYANG M, WELSH W J, GEORGOPOULOS P. Gaussian mixture clustering and imputation of microarray data [J]. Bioinformatics, 2004, 20(6): 917-923.
DING Y, ROSS A. A comparison of imputation methods for handling missing scores in biometric fusion [J]. Pattern Recognition, 2012, 45(3): 919-933.
KANG P. Locally linear reconstruction based missing value imputation for supervised learning [J]. Neurocomputing, 2013, 118(11): 65-78.
张孙力, 杨慧中. 基于改进的K近缺失数据补全 [J]. 计算机与应用化学, 2015, 32(12): 1499-1503.
ZHANG Sunli, YANG Huizhong. Missing data completion based on an improved K-neighbor algorithm [J]. Computers and Applied Chemistry, 2015, 32(12): 1499-1503.
LIU Z G, PAN Q, DEZERT J, et al. Adaptive imputation of missing values for incomplete pattern classification [J]. Pattern Recognition, 2016, 52(C): 85-95.
MUJA M, LOWE D G. Scalable nearest neighbor algorithms for high dimensional data [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2014, 36(11): 2227-2240.
LV Q, JOSEPHSON W, WANG Z, et al. Multi-probe LSH: efficient indexing for high-dimensional similarity search [C]∥2007 33rd International Conference on Very Large Data Bases. San Jose, CA, USA: VLDB Endowment, 2007: 950-961.
WAGAMAN A. Efficient k-neighbor algorithm [J]. Computers and Applied Chemistry, 2015, 32(12): 1499-1503.
LIU Z G, PAN Q, DEZERT J, et al. Adaptive imputation of missing values for incomplete pattern classification [J]. Pattern Recognition, 2016, 52(C): 85-95.
MUJA M, LOWE D G. Scalable nearest neighbor algorithms for high dimensional data [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2014, 36(11): 2227-2240.
LV Q, JOSEPHSON W, WANG Z, et al. Multi-probe LSH: efficient indexing for high-dimensional similarity search [C]∥2007 33rd International Conference on Very Large Data Bases. San Jose, CA, USA: VLDB Endowment, 2007: 950-961.
WAGAMAN A. Efficient graph construction for graphs on variables [J]. Statistical Analysis Data Mining, 2013, 6(6): 443-455.
ALVARO B, DORRONSORO J R. Momentum sequential minimal optimization: an accelerated method for support vector machine training [C]∥International Joint Conference on Neural Networks. Piscataway, NJ, USA: IEEE, 2011: 370-377.
BOYD S, VANDENBERGHE L. Convex optimization [J]. IEEE Transactions on Automatic Control, 2006, 51(11): 1859-1876.
王博,张为,刘艳艳,等.采用机器学习的火焰前景提取算法.2017,51(8):26-32.[doi:10.7652/xjtuxb201708005]
韩江涛,张英杰,李程辉,等.面向光学三维测量的非均匀条纹生成方法.2017,51(8):12-18.[doi:10.7652/xjtuxb201708 003]
徐思雨,蔡佳妮,祝继华,等.自适应多位编码量化的哈希图像检索方法.2017,51(8):19-25.[doi:10.7652/xjtuxb201708 004]
翟夕阳,王晓丹,李睿,等.采用多类代价指数损失函数的代价敏感AdaBoost算法.2017,51(8):33-39.[doi:10.7652/xjtuxb201708006]
付学中,方宗德,关亚彬,等.采用NSGA-Ⅱ算法的面齿轮副小轮拓扑修形多目标优化.2017,51(7):98-104.[doi:10.7652/xjtuxb201707015]
高智勇,董荣光,高建民,等.采用聚类特征的基本概率分配生成方法及应用.2016,50(10):8-14.[doi:10.7652/xjtuxb 201610002]
0
浏览量
5
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621