西安交通大学智能网络与网络安全教育部重点实验室,西安,710049
网络首发:2020-12-10,
纸质出版:2020
移动端阅览
谢雨杰, 杜友田, 张潇. 面向新概念学习的图像描述生成模型[J]. 西安交通大学学报, 2020,54(12):37-44.
An Image Description Generation Model for the Learning of Novel Concepts[J]. 2020, 54(12): 37-44.
谢雨杰, 杜友田, 张潇. 面向新概念学习的图像描述生成模型[J]. 西安交通大学学报, 2020,54(12):37-44. DOI: 10.7652/xjtuxb202012005.
An Image Description Generation Model for the Learning of Novel Concepts[J]. 2020, 54(12): 37-44. DOI: 10.7652/xjtuxb202012005.
为了更好地在图像描述生成任务中对新概念进行学习和预测
在编码-解码框架下提出了一种新的面向新概念学习的图像描述生成模型(Att-DCC)。该模型引入了带有空间注意力机制的卷积神经网络
将全局视觉特征、语义标签和经空间注意力作用后的视觉信息进行了较好的融合; 此外
引入自适应注意力机制多模态层
将语义相近的概念学习结果迁移至新概念
降低训练过程的复杂程度并提升学习性能。采用Att-DCC模型在MSCOCO2014数据集上针对2批(分别为8和6个)共14个新概念进行了测试和分析
结果表明:充分的多模态融合方式和多种注意力机制对于提升学习效果有显著效果; Att-DCC模型在F
1
值上取得了42.56%和42.14%的平均结果
总体上取得了比具有代表性的NOC模型和DCC模型更准确的预测结果。
A new image description generation model with deep compositional captioner based attention mechanism(named Att-DCC model)based on the encoder-decoder framework is proposed to better understand the “novel concepts” that do not appear in the training set for the task of image description generation. A convolutional neural network with spatial attention mechanism is introduced into the model
and global visual features
semantic labels and visual information after spatial attention are well integrated. In addition
a multi-modal layer with an adaptive attention mechanism is introduced to transfer the learning results to novel concepts
which reduces the complexity of training process and improves learning performance. This work extends the existing multi-modal fusion and
introduces multiple attention mechanisms to enhance the performance of novel concept learning. Att-DCC model is used to test and analyze 14 novel concepts in two batches on MSCOCO2014 data set. The experimental results show that the proposed model achieves average results of 42.56% and 42.14% for two batches of novel concepts(8 and 6
respectively)in terms of the F
1
score
and in general
the prediction results are more accurate than those of the representative NOC model and DCC model.
WANG W, YANG Xiaoan, OOI B C, et al. Effective deep learning-based multi-modal retrieval [J]. VLDB Journal, 2016, 25(1): 79-101.
FARHADI A, HEJRATI M, SADEGHI M A, et al. Every picture tells a story: generating sentences from images [C]∥11th European Conference on Computer Vision. Berlin, Germany: Springer Verlag, 2010: 15-29.
LI S, KULKARNI G, BERG T L, et al. Composing simple image descriptions using web-scale n-grams [C]∥Proceedings of the Fifteenth Conference on Computational Natural Language Learning. New York, USA: Association for Computational Linguistics, 2011: 220-228.
FANG Hao, GUPTA S, IANDOLA F, et al. From captions to visual concepts and back [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2015: 1473-1482.
XU K, BA J L, KIROS R, et al. Show, attend and tell: Neural image caption generation with visual attention [C]∥32nd International Conference on Machine Learning. Pitsburg, PA, USA: CMU, 2015: 2048-2057.
BAHDANAU D, CHO K, BENGIO Y. Neural machine translation by jointly learning to align and translate [C]∥3rd International Conference o Learning Representations. Montreal, Canada: ICLR, 2015: 149801.
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need [C]∥31st Annual Conference on Neural Information Processing Systems. Vancouver, Canada: Neural Information Processing Systems Foundation, 2017: 5999-6009.
CHEN L, ZHANG H, XIAO J, et al. Sca-cnn: spatial and channel-wise attention in convolutional networks for image captioning [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5659-5667.
LU J, XIONG C, PARIKH D, et al. Knowing when to look: adaptive attention via a visual sentinel for image captioning [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 3242-3250.
周治平, 张威. 结合视觉属性注意力和残差连接的图像描述生成模型 [J]. 计算机辅助设计与图形学学报, 2018, 30(8): 1536-1542, 1553.
ZHOU Zhiping, ZHANG Wei. An image caption generation model based on visual concept attention and residual connection [J]. Journal of Computer-Aided Design Computer Graphics, 2018, 30(8): 1536-1542, 1553.
汤鹏杰, 谭云兰, 李金忠. 融合图像场景及物体先验知识的图像描述生成模型 [J]. 中国图象图形学报, 2017, 22(9): 1251-1260.
TANG Pengjie, TAN Yunlan, LI Jinzhong. Image description based on the fusion of scene and object category prior knowledge [J]. Journal of Image and Graphics, 2017, 22(9): 1251-1260.
ANNE HENDRICKS L, VENUGOPALAN S, ROHRBACH M, et al. Deep compositional captioning: describing novel object categories without paired training data [C]∥Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2016: 7780377.
YAO Ting, PAN Yingwei, LI Yehao, et al. Incorporating copying mechanism in image captioning for learning novel objects [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5263-5271.
VENUGOPALAN S, ANNE HENDRICKS L, ROHRBACH M, et al. Captioning images with diverse objects [C]∥Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 1170-1178.
WU Yu, ZHU Linchao, JIANG Lu, et al. Decoupled novel object captioner [C]∥Proceedings of the 2018 ACM Multimedia Conference. New York, USA: ACM, 2018: 1029-1037.
FENG Q, WU Y, FAN H, et al. Cascaded revision network for novel object captioning [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2020,30(1): 3413-3421.
MIKOLOV T, CHEN K, CORRADO G, et al. Efficient estimation of word representations in vector space [C]∥1st International Conference of Learning Representation. Montreal, Canada: ICLR, 2013: 149796.
RUSSAKOVSKY O, DENG J, SU H, et al. Imagenet large scale visual recognition challenge [J]. International Journal of Computer Vision, 2015, 115(3): 211-252.
DONAHUE J, HENDRICKS L A, GUADARRAMA S, et al. Long-term recurrent convolutional networks for visual recognition and description [C]∥ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2015, 2625-2634.
PAPINENI K, ROUKOS S, WARD T, et al. BLEU: a method for automatic evaluation of machine translation [EB/OL]. [2020-02-18]. https:∥www.aclweb. org/anthology/P02.1040.pdf.
IBANERJEE S, LAVIE A. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments [EB/OL].[2020-02-20]. https:∥www.aclweb.org/anthology/W05-0909.pdf.
杨楠, 南琳, 张丁一, 等. 基于深度学习的图像描述研究 [J]. 红外与激光工程, 2018, 47(2): 9-16.
YANG Nan, NAN Lin, ZHANG Dingyi, et al. Research on image interpretation based on deep learning [J]. Infrared and Laser Engineering, 2018, 47(2): 9-16.
黄文君,李杰,齐春.低秩与字典表达分解的浓雾霾场景图像去雾算法.2020,54(4):118-125.[doi:10.7652/xjtuxb202004 015]
毛远宏,贺占庄,马钟,等.采用类内迁移学习的红外/可见光异源图像匹配.2020,54(1):49-55.[doi:10.7652/xjtuxb 202001007]
贾晓芬,郭永存,柴华荣,等.深立井井壁图像的卷积神经网络去噪方法.2019,53(6):117-124.[doi:10.7652/xjtuxb2019 06016]
崔云博,杜友田,王航.面向图像的有效目标区域提取方法.2019,53(5):52-57.[doi:10.7652/xjtuxb201905008]
肖满生,肖哲,万烂军.多特征融合的图像格贴近度匹配方法.2019,53(4):115-121.[doi:10.7652/xjtuxb201904017]
孟月波,刘光辉,徐胜军,等.一种具有边缘保持的多尺度马尔可夫随机场模型图像分割方法.2019,53(3):56-65.[doi:10.7652/xjtuxb201903009]
时璇,许林松,李晨,等.联合加权聚合深度卷积特征的图像检索方法.2019,53(2):128-135.[doi:10.7652/xjtuxb2019 02017]
王浩,梁煜,张为.一种均匀化稀疏表示的图像压缩感知算法.2019,53(2):136-141.[doi:10.7652/xjtuxb201902018]
邢超,张晓明,赵亚琳,等.利用图像信息熵差检测炉渣碳质量分数.2018,52(8):49-53.[doi:10.7652/xjtuxb201808008]
宋长明,王赟.融合低秩和稀疏表示的图像超分辨率重建算法.2018,52(7):18-24.[doi:10.7652/xjtuxb201807003]
张淑芳,丁文鑫,韩泽欣,等.采用主成分分析与梯度金字塔的高动态范围图像生成方法.2018,52(4):150-157.[doi:10.7652/xjtuxb201804022]
徐思雨,蔡佳妮,祝继华,等.自适应多位编码量化的哈希图像检索方法.2017,51(8):19-25.[doi:10.7652/xjtuxb201708 004]
李莉,冯林,吴俊,等.一种三结构描述子的图像检索方法.2017,51(6):86-91.[doi:10.7652/xjtuxb201706014]
张胜杰,查宇飞,李运强,等.一种利用局部结构信息的加权哈希图像检索算法.2016,50(10):78-85.[doi:10.7652/xjtuxb201610012]
黄晓冬,孙亮,刘胜蓝.一种判别极端学习的相关反馈图像检索方法.2016,50(8):96-102.[doi:10.7652/xjtuxb201608 016]
周远,周玉生,刘权,等.一种适用于图像拼接的DSIFT算法研究.2015,49(9):84-90.[doi:10.7652/xjtuxb201509015]
毛彦斌,张选平,杨晓刚.伪DNA密码图像加密算法研究.2015,49(9):91-98.[doi:10.7652/xjtuxb201509016]
刘凯,张立民,孙永威,等.利用深度玻尔兹曼机与典型相关分析的自动图像标注算法.2015,49(6):33-38.[doi:10.7652/xjtuxb201506006]
符均,牟轩沁,季文博.亮色分离的饱和图像校正方法.2014,48(10):101-107.[doi:10.7652/xjtuxb201410016]
岳桂华,滕奇志,何小海,等.岩心三维图像修复算法.2014,48(9):37-42.[doi:10.7652/xjtuxb201409007]
0
浏览量
4
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621