A new image description generation model with deep compositional captioner based attention mechanism(named Att-DCC model)based on the encoder-decoder framework is proposed to better understand the “novel concepts” that do not appear in the training set for the task of image description generation. A convolutional neural network with spatial attention mechanism is introduced into the model
and global visual features
semantic labels and visual information after spatial attention are well integrated. In addition
a multi-modal layer with an adaptive attention mechanism is introduced to transfer the learning results to novel concepts
which reduces the complexity of training process and improves learning performance. This work extends the existing multi-modal fusion and
introduces multiple attention mechanisms to enhance the performance of novel concept learning. Att-DCC model is used to test and analyze 14 novel concepts in two batches on MSCOCO2014 data set. The experimental results show that the proposed model achieves average results of 42.56% and 42.14% for two batches of novel concepts(8 and 6
respectively)in terms of the F
1
score
and in general
the prediction results are more accurate than those of the representative NOC model and DCC model.
关键词
Keywords
references
WANG W, YANG Xiaoan, OOI B C, et al. Effective deep learning-based multi-modal retrieval [J]. VLDB Journal, 2016, 25(1): 79-101.
FARHADI A, HEJRATI M, SADEGHI M A, et al. Every picture tells a story: generating sentences from images [C]∥11th European Conference on Computer Vision. Berlin, Germany: Springer Verlag, 2010: 15-29.
LI S, KULKARNI G, BERG T L, et al. Composing simple image descriptions using web-scale n-grams [C]∥Proceedings of the Fifteenth Conference on Computational Natural Language Learning. New York, USA: Association for Computational Linguistics, 2011: 220-228.
FANG Hao, GUPTA S, IANDOLA F, et al. From captions to visual concepts and back [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2015: 1473-1482.
XU K, BA J L, KIROS R, et al. Show, attend and tell: Neural image caption generation with visual attention [C]∥32nd International Conference on Machine Learning. Pitsburg, PA, USA: CMU, 2015: 2048-2057.
BAHDANAU D, CHO K, BENGIO Y. Neural machine translation by jointly learning to align and translate [C]∥3rd International Conference o Learning Representations. Montreal, Canada: ICLR, 2015: 149801.
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need [C]∥31st Annual Conference on Neural Information Processing Systems. Vancouver, Canada: Neural Information Processing Systems Foundation, 2017: 5999-6009.
CHEN L, ZHANG H, XIAO J, et al. Sca-cnn: spatial and channel-wise attention in convolutional networks for image captioning [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5659-5667.
LU J, XIONG C, PARIKH D, et al. Knowing when to look: adaptive attention via a visual sentinel for image captioning [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 3242-3250.
ZHOU Zhiping, ZHANG Wei. An image caption generation model based on visual concept attention and residual connection [J]. Journal of Computer-Aided Design Computer Graphics, 2018, 30(8): 1536-1542, 1553.
TANG Pengjie, TAN Yunlan, LI Jinzhong. Image description based on the fusion of scene and object category prior knowledge [J]. Journal of Image and Graphics, 2017, 22(9): 1251-1260.
ANNE HENDRICKS L, VENUGOPALAN S, ROHRBACH M, et al. Deep compositional captioning: describing novel object categories without paired training data [C]∥Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2016: 7780377.
YAO Ting, PAN Yingwei, LI Yehao, et al. Incorporating copying mechanism in image captioning for learning novel objects [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5263-5271.
VENUGOPALAN S, ANNE HENDRICKS L, ROHRBACH M, et al. Captioning images with diverse objects [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 1170-1178.
WU Yu, ZHU Linchao, JIANG Lu, et al. Decoupled novel object captioner [C]∥Proceedings of the 2018 ACM Multimedia Conference. New York, USA: ACM, 2018: 1029-1037.
FENG Q, WU Y, FAN H, et al. Cascaded revision network for novel object captioning [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2020,30(1): 3413-3421.
MIKOLOV T, CHEN K, CORRADO G, et al. Efficient estimation of word representations in vector space [C]∥1st International Conference of Learning Representation. Montreal, Canada: ICLR, 2013: 149796.
RUSSAKOVSKY O, DENG J, SU H, et al. Imagenet large scale visual recognition challenge [J]. International Journal of Computer Vision, 2015, 115(3): 211-252.
DONAHUE J, HENDRICKS L A, GUADARRAMA S, et al. Long-term recurrent convolutional networks for visual recognition and description [C]∥ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2015, 2625-2634.
PAPINENI K, ROUKOS S, WARD T, et al. BLEU: a method for automatic evaluation of machine translation [EB/OL]. [2020-02-18]. https:∥www.aclweb. org/anthology/P02.1040.pdf.
IBANERJEE S, LAVIE A. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments [EB/OL].[2020-02-20]. https:∥www.aclweb.org/anthology/W05-0909.pdf.
YANG Nan, NAN Lin, ZHANG Dingyi, et al. Research on image interpretation based on deep learning [J]. Infrared and Laser Engineering, 2018, 47(2): 9-16.