

浏览全部资源
扫码关注微信
1.新疆大学计算机科学与技术学院, 830046,乌鲁木齐
2.西安交通大学电子与信息学部, 710049,西安
Received:09 November 2024,
Online First:07 February 2025,
Published:10 June 2025
移动端阅览
YANG Xiangyan, LIANG Huihui, CHEN Xi, et al. Audio-Driven Talking Face Generation with Multi-Scale Visual Enhancement[J]. Journal of Xi’an Jiaotong University, 2025, 59(6): 167-176.
YANG Xiangyan, LIANG Huihui, CHEN Xi, et al. Audio-Driven Talking Face Generation with Multi-Scale Visual Enhancement[J]. Journal of Xi’an Jiaotong University, 2025, 59(6): 167-176. DOI: 10.7652/xjtuxb202506017.
针对现有语音驱动人脸生成方法在生成人脸视频清晰度和真实感方面的不足,设计了一种多尺度视觉增强的端到端人脸视频生成模型——VisClearTalk,并提出包含视觉增强模块的人脸解码器。首先,采用人脸编码器处理包含随机参考帧和下半张脸被遮挡的先验帧,从中提取面部特征;接着,采用语音编码器从音频中提取特征,以指导面部内容生成;随后,采用人脸解码器融合上述特征,并通过卷积模块初步重建面部图像;最后,采用视觉增强模块将多尺度卷积和残差融合,进一步增强人脸下半部分区域的细节和边缘信息,从而提升生成人脸视频的视觉质量。使用公开唇语识别数据集对VisClearTalk模型进行实验验证,结果表明:通过引入视觉增强模块,该模型有效提升了面部视觉内容的细腻程度和真实感,能够生成清晰自然的人脸视频;在性能指标方面,峰值信噪比达到34.349 dB,结构相似度达到0.933,可学习感知图像块相似度降低至0.040。研究结果可为当前人脸生成需求提供一种解决方案。
To address the limitations of existing audio-driven talking face generation methods in terms of video clarity and realism
an end-to-end talking face generation method called VisClearTalk which incorporates multi-scale visual enhancement is proposed in this paper
and a face decoder with a visual enhancement module is proposed. First
the face encoder processed a random reference frame and a prior frame with the lower half of the face occluded to extract facial features. Simultaneously
the audio encoder extracted features from the audio to guide facial content generation. Subsequently
the face decoder integrated these features and performed an initial reconstruction of facial images through convolutional modules. Finally
the visual enhancement module employed multi-scale convolution and residual fusion to further enhance the details and edge information of the lower face region
improving the visual quality of the generated talking face videos. The VisClearTalk model was experimentally validated using public lip-reading datasets
with both quantitative and qualitative results demonstrating that the introduction of the visual enhancement module effectively improves the fineness and realism of facial visual content
enabling the generation of clear and natural talking face videos. In terms of performance metrics
the peak signal-to-noise ratio reached 34.349 dB
structural similarity reached 0.933
and learnable perceptual image patch similarity was reduced to 0.040. The VisClearTalk model offers a viable solution for current talking face videos generation needs.
YU Lingyun , YU Jun , LI Mengyan , et al . Multimodal inputs driven talking face generation with spatial-temporal dependency [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2021 , 31 ( 1 ): 203 - 216 .
TOSHPULATOV M , LEE W , LEE S . Talking human face generation: a survey [J ] . Expert Systems with Applications , 2023 , 219 : 119678 .
张哲 , 齐春 , 张钊强 , 等 . 人脸超分辨率重建中投影空间的选择方法 [J ] . 西安交通大学学报 , 2018 , 52 ( 8 ): 43 - 48 .
ZHANG Zhe , QI Chun , ZHANG Zhaoqiang , et al . A selection method of projection space for face super-resolution reconstruction [J ] . Journal of Xi’an Jiaotong University , 2018 , 52 ( 8 ): 43 - 48 .
张云佐 , 郭亚宁 , 李文博 . 融合时空切片和双注意力机制的视频摘要方法 [J ] . 西安交通大学学报 , 2022 , 56 ( 12 ): 127 - 135 .
ZHANG Yunzuo , GUO Yaning , LI Wenbo . Video summarization method based on spatiotemporal slice and dual attention mechanism [J ] . Journal of Xi’an Jiaotong University , 2022 , 56 ( 12 ): 127 - 135 .
ZHONG Weizhi , FANG Chaowei , CAI Yinqi , et al . Identity-preserving talking face generation with landmark and appearance priors [C ] // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2023 : 9729 - 9738 .
Chatziagapi A , Athar S , JAIN A , et al . LipNeRF: what is the right feature space to lip-sync a NeRF? [C ] // 2023 IEEE 17th International Conference on Automatic Face and Gesture Recognition (FG) . Piscataway, NJ, USA : IEEE , 2023 : 1 - 8 .
ZHANG Wenxuan , CUN Xiaodong , WANG Xuan , et al . SadTalker: learning realistic 3D motion coefficients for stylized audio-driven single image talking face animation [C ] // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2023 : 8652 - 8661 .
孙威 . 中文文本驱动的人脸说话视频生成方法研究与实现 [D ] . 南京 : 东南大学 , 2022 .
ESKIMEZ S E , ZHANG You , DUAN Zhiyao . Speech driven talking face generation from a single image and an emotion condition [J ] . IEEE Transactions on Multimedia , 2022 , 24 : 3480 - 3490 .
YE Zhenhui , JIANG Ziyue , REN Yi , et al . GeneFace: generalized and high-fidelity audio-driven 3D talking face synthesis [EB/OL ] . ( 2023-01-31 ) [ 2024-05-15 ] . https://arxiv.org/abs/2301.13430 https://arxiv.org/abs/2301.13430 .
YE Zhenhui , HE Jinzheng , JIANG Ziyue , et al . GeneFace++: generalized and stable real-time audio-driven 3D talking face generation [EB/OL ] . ( 2023-05-01 ) [ 2024-06-20 ] . https://arxiv.org/abs/2305.00787 https://arxiv.org/abs/2305.00787 .
CHEN Lele , MADDOX R K , DUAN Zhiyao , et al . Hierarchical cross-modal talking face generation with dynamic pixel-wise loss [C ] // 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2019 : 7824 - 7833 .
YE Zipeng , XIA Mengfei , YI Ran , et al . Audio-driven talking face video generation with dynamic convolution kernels [J ] . IEEE Transactions on Multimedia , 2023 , 25 : 2033 - 2046 .
PARK S J , KIM M , HONG J , et al . SyncTalkFace: talking face generation with precise lip-syncing via audio-lip memory [C ] // Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence and Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence and the Twelveth Symposium on Educational Advances in Artificial Intelligence . Palo Alto, CA, USA : AAAI Press , 2022 : 2062 - 2070 .
侯冰莹 . 基于生成式对抗网络的语音驱动人脸生成的研究 [D ] . 武汉 : 华中师范大学 , 2023 .
SUWAJANAKORN S , SEITZ S M , KEMELMACHER-SHLIZERMAN I . Synthesizing Obama: learning lip sync from audio [J ] . ACM Transactions on Graphics , 2017 , 36 ( 4 ): 95 .
ZHENG Ruobing , ZHU Zhou , SONG Bo , et al . A neural lip-sync framework for synthesizing photorealistic virtual news anchors [C ] // 2020 25th International Conference on Pattern Recognition (ICPR) . Piscataway, NJ, USA : IEEE , 2021 : 5286 - 5293 .
PRAJWAL K R , MUKHOPADHYAY R , NAMBOODIRI V P , et al . A lip sync expert is all you need for speech to lip generation in the wild [C ] // Proceedings of the 28th ACM International Conference on Multimedia . New York, USA : ACM , 2020 : 484 - 492 .
ZHANG Zhimeng , HU Zhipeng , DENG Wenjin , et al . Dinet: deformation inpainting network for realistic face visually dubbing on high resolution video [C ] // AAAI Conference on Artificial Intelligence . Palo Alto, CA, USA : AAAI Press , 2023 , 37 : 3543 - 3551 .
CHEN Yaosen , YAO Yu , LI Zhiqiang , et al . HyperLips: hyper control lips with high resolution decoder for talking face generation [J ] . Applied Intelligence , 2024 , 55 ( 2 ): 145 .
NEWELL A , YANG Kaiyu , DENG Jia . Stacked hourglass networks for human pose estimation [C ] // Computer Vision-ECCV 2016 . Cham : Springer International Publishing , 2016 : 483 - 499 .
SIMONYAN K , ZISSERMAN A . Very deep convolutional networks for large-scale image recognition [EB/OL ] . ( 2015-04-10 ) [ 2024-10-12 ] . https://arxiv.org/abs/1409.1556 https://arxiv.org/abs/1409.1556 .
ZHANG R , ISOLA P , EFROS A A , et al . The unreasonable effectiveness of deep features as a perceptual metric [C ] // 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway, NJ, USA : IEEE , 2018 : 586 - 595 .
KINGMA D P , BA J . Adam: a method for stochastic optimization [EB/OL ] . ( 2017-01-30 ) [ 2024-10-20 ] . https://arxiv.org/abs/1412.6980 https://arxiv.org/abs/1412.6980 .
ZHAO Ya , XU Rui , SONG Mingli . A cascade sequence-to-sequence model for Chinese mandarin lip reading [C ] // Proceedings of the 1st ACM International Conference on Multimedia in Asia . New York, USA : ACM , 2020 : 32 .
HORÉ A , ZIOU D . Image quality metrics: PSNR vs. SSIM [C ] // 2010 20th International Conference on Pattern Recognition . Piscataway, NJ, USA : IEEE , 2010 : 2366 - 2369 .
WANG Zhou , BOVIK A C , SHEIKH H R , et al . Image quality assessment: from error visibility to structural similarity [J ] . IEEE Transactions on Image Processing , 2004 , 13 ( 4 ): 600 - 612 .
AFOURAS T , CHUNG J S , SENIOR A , et al . Deep audio-visual speech recognition [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022 , 44 ( 12 ): 8717 - 8727 .
0
Views
26
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621