1.福建江夏学院电子信息科学学院, 350108,福州
2.西安交通大学电子与信息学部, 710049,西安
3.厦门理工学院电气工程与自动化学院, 361024,福建厦门
4.兰州理工大学电气工程与信息工程学院, 730050,兰州
谢玉枚(1986—),女,副教授;
关翔锋,男,教授。
收稿:2024-09-30,
网络首发:2025-02-26,
纸质出版:2025-06-10
移动端阅览
谢玉枚, 蔡远利, 高海燕, 等. 融合U-net网络的纯卷积视频预测模型[J]. 西安交通大学学报, 2025,59(6):112-121.
XIE Yumei, CAI Yuanli, GAO Haiyan, et al. A Pure Convolutional Model Fused with U-net Network for Video Prediction[J]. Journal of Xi’an Jiaotong University, 2025, 59(6): 112-121.
谢玉枚, 蔡远利, 高海燕, 等. 融合U-net网络的纯卷积视频预测模型[J]. 西安交通大学学报, 2025,59(6):112-121. DOI: 10.7652/xjtuxb202506012.
XIE Yumei, CAI Yuanli, GAO Haiyan, et al. A Pure Convolutional Model Fused with U-net Network for Video Prediction[J]. Journal of Xi’an Jiaotong University, 2025, 59(6): 112-121. DOI: 10.7652/xjtuxb202506012.
为了解决基于深度学习视频预测中存在的时空特征提取不充分以及图像细节保留不足的问题,运用简单视频预测网络模型SimVP给出的Inception单元,提出了一种融合U-net网络的纯卷积视频预测模型(CUnet)。CUnet模型由3个核心模块组成:首先,Cell模块采用2D卷积层来提取空间特征,并将这些特征输入至多个Inception单元捕获时空特性;其次,DeCell模块通过Inception单元捕获时空特征,并借助2D反卷积层进行上采样操作,恢复图像原始尺寸;最后,引入U-net作为主干网络,将Cell模块和DeCell模块有机整合,有效保留了图像的细节信息,实现了高质量的图像重建。实验结果表明:在TaxiBJ数据集上,与当前表现最佳的时间注意力单元网络模型TAU相比,CUnet模型的预测精度提高了5.23%;在Human3.6M数据集上,与当前表现最佳的快速傅里叶Inception网络模型FFINet相比,CUnet模型的预测精度提高了12.88%。CUnet模型具有优秀的预测能力,可为纯卷积神经网络模型在视频预测领域的应用提供有益探索。
To address the issues of insufficient spatiotemporal feature extraction and inadequate image detail preservation in deep learning-based video prediction
a pure convolutional video prediction model (CUnet) fused with the U-net network
using the Inception unit from the SimVP model
is proposed. CUnet model consists of 3 core modules. Firstly
the Cell module uses 2D convolutional layers to extract spatial features and feeds these features into multiple Inception units to capture spatiotemporal features. Secondly
the DeCell module captures spatiotemporal features through Inception units and performs upsampling operations using 2D deconvolutional layers to restore the original image size. Finally
U-net is introduced as the backbone network to organically integrate the Cell module and the DeCell module
effectively preserving the detailed information of the image and achieving high-quality image reconstruction. The experimental results showed that on the TaxiBJ dataset
compared with the currently best-performing TAU model
the prediction accuracy of the CUnet model had increased by 5.23%. On the Human3.6M dataset
compared with the currently best-performing FFINet model
the prediction accuracy of the CUnet model had increased by 12.88%. The CUnet model demonstrates exceptional predictive capabilities
offering valuable insights for the application of pure convolutional neural networks in the field of video prediction.
李善梅 , 宋思霓 , 王红勇 , 等 . 基于多模态时空特征融合的终端区交通拥堵精细化预测 [J/OL ] . 北京航空航天大学学报 . ( 2024-09-05 ) [ 2024-09-18 ] . https://doi.org/10.13700/j.bh.1001-5965.2024.0557 https://doi.org/10.13700/j.bh.1001-5965.2024.0557 .
LI Shanmei , SONG Sini , WANG Hongyong , et al . Refined prediction of terminal area traffic congestion based on multimodal spatiotemporal feature fusion [J/OL ] . Journal of Beijing University of Aeronautics and Astronautics . ( 2024-09-05 ) [ 2024-09-18 ] . https://doi.org/10.13700/j.bh.1001-5965.2024.0557 https://doi.org/10.13700/j.bh.1001-5965.2024.0557 .
NING Shuliang , LAN Mengcheng , LI Yanran , et al . MIMO is all you need: a strong multi-in-multi-out baseline for video prediction [C ] // Proceedings of the AAAI Conference on Artificial Intelligence . Palo Alto, CA, USA : AAAI Press , 2023 : 1975 - 1983 .
吴宇轩 , 虞慧群 , 范贵生 . 基于误差补偿的多模态协同交通流预测模型 [J ] . 电子学报 , 2024 , 52 ( 8 ): 2878 - 2890 .
WU Yuxuan , YU Huiqun , FAN Guisheng . Multimodal cooperative traffic flow prediction model based on error compensation [J ] . Acta Electronica Sinica , 2024 , 52 ( 8 ): 2878 - 2890 .
CHENG Jinguo , LI Ke , LIANG Yuxuan , et al . Rethinking urban mobility prediction: a super-multivariate time series forecasting approach [EB/OL ] . ( 2023-12-04 ) [ 2024-08-10 ] . https://arxiv.org/abs/2312.01699 https://arxiv.org/abs/2312.01699 .
晏婕 . 基于深度学习的视频帧序列预测算法研究 [D ] . 长春 : 吉林大学 , 2023 .
TANG Yujin , DONG Peijie , TANG Zhenheng , et al . VMRNN: Integrating vision mamba and LSTM for efficient and accurate spatiotemporal forecasting [EB/OL ] . ( 2024-06-29 ) [ 2024-08-12 ] . https://arxiv.org/abs/2403.16536 https://arxiv.org/abs/2403.16536 .
孔玮 , 刘云 , 李辉 , 等 . 基于深度学习的行人轨迹预测方法综述 [J ] . 控制与决策 , 2021 , 36 ( 12 ): 2841 - 2850 .
KONG Wei , LIU Yun , LI Hui , et al . Survey of pedestrian trajectory prediction methods based on deep learning [J ] . Control and Decision , 2021 , 36 ( 12 ): 2841 - 2850 .
潘敏婷 , 王韫博 , 朱祥明 , 等 . 基于无标签视频数据的深度预测学习方法综述 [J ] . 电子学报 , 2022 , 50 ( 4 ): 869 - 886 .
PAN Minting , WANG Yunbo , ZHU Xiangming , et al . A survey on deep predictive learning based on unlabeled videos [J ] . Acta Electronica Sinica , 2022 , 50 ( 4 ): 869 - 886 .
SHI Xingjian , CHEN Zhourong , WANG Hao , et al . Convolutional LSTM network: a machine learning approach for precipitation nowcasting [C ] // Proceedings of the 28th International Conference on Neural Information Processing Systems . Cambridge, MA, USA : MIT Press , 2015 : 802 - 810 .
WANG Yunbo , LONG Mingsheng , WANG Jianmin , et al . PredRNN: recurrent neural networks for predictive learning using spatiotemporal LSTMs [C ] // Proceedings of the 31st International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2017 : 879 - 888 .
WANG Yunbo , GAO Zhifeng , LONG Mingsheng , et al . PredRNN++: towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning [C ] // Proceedings of the 35th International Conference on Machine Learning . Chia Laguna Resort, Sardinia, Italy : PMLR , 2018 : 5123 - 5132 .
WANG Yunbo , ZHANG Ananjin , ZHU Hongyu , et al . Memory in memory: a predictive neural network for learning higher-order non-stationarity from spatiotemporal dynamics [C ] // 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2019 : 9146 - 9154 .
WANG Yunbo , JIANG Lu , YANG M H , et al . Eidetic 3D LSTM: a model for video prediction and beyond [C ] // 7th International Conference on Learning Representations . New York, USA : ICLR , 2019 : 1 - 14 .
YU Wei , LU Yichao , EASTERBROOK S , et al . Efficient and information-preserving future frame prediction and beyond [C ] // International Conference on Learning Representations . New York, USA : ICLR , 2020 : 1 - 14 .
LE GUEN V , THOME N . Disentangling physical dynamics from unknown factors for unsupervised video prediction [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2020 : 11471 - 11481 .
CHANG Zheng , ZHANG Xinfeng , WANG Shanshe , et al . MAU: a motion-aware unit for video prediction and beyond [C ] // Proceedings of the 35th International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2021 : 26950 - 26962 .
GAO Zhangyang , TAN Cheng , WU Lirong , et al . SimVP: simpler yet better video prediction [C ] // 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2022 : 3160 - 3170 .
LI Ping , ZHANG Chenhan , XU Xianghua . Fast Fourier inception networks for occluded video prediction [J ] . IEEE Transactions on Multimedia , 2024 , 26 : 3418 - 3429 .
TAN Cheng , GAO Zhangyang , WU Lirong , et al . Temporal attention unit: towards efficient spatiotemporal predictive learning [C ] // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2023 : 18770 - 18782 .
WEN Shiping , LIU Weiwei , YANG Yin , et al . Generating realistic videos from keyframes with concatenated GANs [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2019 , 29 ( 8 ): 2337 - 2348 .
KWON Y H , PARK M G . Predicting future frames using retrospective cycle GAN [C ] // 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2019 : 1811 - 1820 .
YU S , TACK J , MO S , et al . Generating videos with dynamics-aware implicit generative adversarial networks [C ] // 10th International Conference on Learning Representations . [S.l. ] : [s.n. ] , 2022 : 1 - 14 .
OPREA S , MARTINEZ-GONZALEZ P , GARCIA-GARCIA A , et al . A review on deep learning techniques for video prediction [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022 , 44 ( 6 ): 2806 - 2826 .
ZHANG Qianqian , FENG Guorui , WU Hanzhou . Surveillance video anomaly detection via non-local U-Net frame prediction [J ] . Multimedia Tools and Applications , 2022 , 81 ( 19 ): 27073 - 27088 .
GAN K Y , CHENG Yutong , TAN H K , et al . Contrastive-regularized U-Net for video anomaly detection [J ] . IEEE Access , 2023 , 11 : 36658 - 36671 .
张杰 , 杨雪 , 龚智龙 , 等 . 双向预测BiP-GAN的行人视频异常事件自动检测 [J/OL ] . 武汉大学学报(信息科学版) . ( 2025-02-11 ) [ 2025-02-18 ] . https://doi.org/10.13203/j.whugis20240259 https://doi.org/10.13203/j.whugis20240259 .
ZHANG Jie , YANG Xue , GONG Zhilong , et al . Bidirectional prediction BiP-GAN pedestrian video anomaly event automatic detection [J/OL ] . Geomatics and Information Science of Wuhan University . ( 2025-02-11 ) [ 2025-02-18 ] . https://doi.org/10.13203/j.whugis20240259 https://doi.org/10.13203/j.whugis20240259 .
WANG Zhaohua , LI Zhenyu , PAN Heng , et al . Large-scale measurements and prediction of DC-WAN traffic [J ] . IEEE Transactions on Parallel and Distributed Systems , 2023 , 34 ( 5 ): 1390 - 1405 .
LEI Shanzhong , SONG Junfang , WANG Tengjiao , et al . Attention U-Net based on multi-scale feature extraction and WSDAN data augmentation for video anomaly detection [J ] . Multimedia Systems , 2024 , 30 ( 3 ): 118 .
TAN Cheng , LI Siyuan , GAO Zhangyang , et al . OpenSTL: a comprehensive benchmark of spatio-temporal predictive learning [C ] // Proceedings of the 37th International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2023 : 69819 - 69831 .
0
浏览量
13
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621