An Automatic Image Annotation Algorithm Using Deep Boltzmann Machine and Canonical Correlation Analysis[J]. 2015, 49(6): 33-38.
DOI:
An Automatic Image Annotation Algorithm Using Deep Boltzmann Machine and Canonical Correlation Analysis[J]. 2015, 49(6): 33-38.DOI: 10.7652/xjtuxb201506006.
An Automatic Image Annotation Algorithm Using Deep Boltzmann Machine and Canonical Correlation Analysis
An automatic image annotation algorithm is proposed based on deep Boltzmann machine and canonical correlation analysis
named DBM-CCA. The algorithm utilizes DBM to transform low-level features of images and labels to sparse high-level abstract concepts
and builds subspace mapping relations by CCA in order to generate labels. The multiple Bernoulli distribution is used to fit labels and the Gaussian distribution is used to fit image features in the process of using DBM to extract high-level features of images and labels. CCA is used to establish relevant connection among image features and labeling words which form canonical variable subspace. High-level text features are calculated based on the Mahalanobis distance between images in canonical variable subspace
and image annotation words are generated by mean-field inference. Experimental results show that the proposed automatic image annotation method significantly outperforms both the traditional MBRM and the SML
and the precision ratio and recall-precision mean ratio are increased by 10% and 5%
respectively
in experiments with Corel5K image dataset.
关键词
Keywords
references
LI Q, GU Y, QIAN X. LCMKL: latent-community and multi-kernel learning based image annotation [C]∥Proceedings of the 22nd ACM International Conference on Information & Knowledge Management. New York, USA: ACM, 2013: 1469-1472.
QIAN X, HUA X S, HOU X. Tag filtering based on similar compatible principle [C]∥Proceedings of IEEE International Conference on Image Processing. Piscataway, NJ, USA: IEEE, 2012: 2349-2352.
QIAN X, HUA X S, TANG Y Y, et al. Social image tagging with diverse semantics [J]. IEEE Transactions on Cybernetics, 2014, 44(12): 2493-2508.
NGIAM J, KHOSLA A, KIM M, et al. Multimodal deep learning [C]∥Proceedings of the 28th International Conference on Machine Learning. New York, USA: ACM, 2011: 689-696.
OUYANG W, CHU X, WANG X. Multi-source deep learning for human pose estimation [C]∥Proceedings of IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2014: 2337-2344.
KIROS R, ZEMEL R, SALAKHUTDINOV R. Multimodal neural language models [J]. Journal of Machine Learning Research, 2014, 32(1): 595-603.
QIU Lida, LIU Tianjian, LIN Nan, et al. Data aggregation in wireless sensor network based on deep learning model [J]. Chinese Journal of Sensors and Actuators, 2014, 27(12): 1704-1709.
SRIVASTAVA N, SALAKHUTDINOV R. Multimodal learning with deep Boltzmann machines [C]∥Proceedings of Advances in Neural Information Processing Systems. Cambridge, MA, USA: MIT, 2012: 2222-2230.
GAO Junfeng, ZHENG Chongxun, WANG Pei. Electromyography artifact removal from electroencephalogram in real-time [J]. Journal of Xi'an Jiaotong University, 2010, 44(4): 114-118.
RASIWASIA N. A new approach to cross-modal multimedia retrieval [C]∥Proceedings of the 18 th ACM International Conference on Multimedia. New York, USA: ACM, 2010: 251-260.
FENG F, WANG X, LI R. Cross-modal retrieval with correspondence autoencoder [C]∥Proceedings of the 22nd ACM International Conference on Multimedia. New York, USA: ACM, 2014: 7-16.
GALEN A, RAMAN A, JEFF B. Deep canonical correlation analysis [J]. Journal of Machine Learning Research, 2013, 28(3): 1247-1255.
SALAKHUTDINOV R, HINTON G E. Deep Boltzmann machines [C]∥Proceedings of International Conference on Artificial Intelligence and Statistics 2009. Brookline, MA, USA: Microtome Publishing, 2009: 448-455.
MAKADIA A, PAVLOVIC V, KUMAR S. Baselines for image annotation [J]. International Journal on Computer Vision, 2010, 90(1): 88-105.