

浏览全部资源
扫码关注微信
1.西北工业大学电子信息学院, 710129,西安
2.皇家墨尔本理工大学工程学院, VIC3001,澳大利亚墨尔本
Received:09 May 2024,
Online First:12 September 2024,
Published:10 February 2025
移动端阅览
ZOU Chengyi, WAN Shuai, ZHU Zhiwei, et al. Cross-Component Prediction for H.266/Versatile Video Coding Based on Lightweight Convolutional Neural Network[J]. Journal of Xi’an Jiaotong University, 2025, 59(2): 180-188.
ZOU Chengyi, WAN Shuai, ZHU Zhiwei, et al. Cross-Component Prediction for H.266/Versatile Video Coding Based on Lightweight Convolutional Neural Network[J]. Journal of Xi’an Jiaotong University, 2025, 59(2): 180-188. DOI: 10.7652/xjtuxb202502018.
为提高新一代通用视频编码标准(H.266/VVC)中色度帧内预测的准确度,提出了采用轻量级卷积神经网络的跨分量预测方法。设计了亮度模块和边界模块,从亮度和色度参考样本中提取特征。设计了注意力模块,构建当前亮度参考样本和边界亮度参考样本之间的空间关系,并应用于边界色度参考样本生成色度预测样本。为降低编解码复杂度,设计网络在二维完成特征融合和预测,优化了现有的同组参数处理不同块大小的训练策略。并且,引入宽度可变卷积,根据不同的块大小调整网络参数。实验结果表明:与H.266/VVC测试模型VTM18.0相比,所提网络在Y(亮度分量)、Cb(蓝色色度分量)、Cr(红色色度分量)上分别实现了0.30%、2.46%、2.25%的码率节省。与其他基于卷积神经网络的跨分量预测方法相比,有效地降低了网络参数和推理复杂度,分别节省了约10%的编码时间和19%的解码时间。
To improve the accuracy of intra chroma prediction in H.266/versatile video coding (VVC)
a cross-component prediction method based on lightweight convolutional neural network was proposed in this paper. The luma module and chroma module were designed to extract features from luma and chroma reference samples
and the attention module was designed to leverage the attention mechanism to construct the spatial correlation between the current luma reference samples and the boundary luma reference samples. Finally
the attention mask was applied to the boundary chroma reference samples to generate chroma prediction value. To reduce the encoding and decoding complexity
the feature fusion and prediction in the network were achieved in two dimensions
the existing training strategy with shared parameters to handle variable block sizes was improved
and slimmable convolutions were introduced to adjust network parameters according to different block sizes. The experimental results show that the proposed algorithm achieved 0.30%/2.46%/2.25% BD-rate reduction on the Y/Cb/Cr component
respectively
compared with the H.266/VVC test model VTM18.0. Compared with other convolutional neural networks-based cross-component prediction methods
the proposed method effectively reduced the network parameters and inference complexity
saving 10% encoding time and 19% decoding time.
SULLIVAN G J , OHM J R , HAN W J , et al . Overview of the high efficiency video coding (HEVC) standard [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2012 , 22 ( 12 ): 1649 - 1668 .
BROSS B , CHEN Jianle , OHM J R , et al . Developments in international video coding standardization after AVC, with an overview of versatile video coding (VVC) [J ] . Proceedings of the IEEE , 2021 , 109 ( 9 ): 1463 - 1493 .
万帅 , 霍俊彦 , 马彦卓 , 等 . 新一代通用视频编码标准H.266/VVC:现状与发展 [J ] . 西安交通大学学报 , 2024 , 58 ( 4 ): 1 - 17 .
WAN Shuai , HUO Junyan , MA Yanzhuo , et al . The new-generation versatile video coding standard H.266/VVC: state-of-the-art and development [J ] . Journal of Xi'an Jiaotong University , 2024 , 58 ( 4 ): 1 - 17 .
万帅 , 霍俊彦 , 马彦卓 , 等 . 新一代通用视频编码H.266/VVC:原理、标准与实现 [M ] . 北京 : 电子工业出版社 , 2022 .
LEE S H , CHO N I . Intra prediction method based on the linear relationship between the channels for YUV 4:2:0 intra coding [C ] // 2009 16th IEEE International Conference on Image Processing (ICIP) . Piscataway, NJ, USA : IEEE , 2009 : 1037 - 1040 .
LI Yue , LI Li , LI Zhu , et al . A hybrid neural network for chroma intra prediction [C ] // 2018 25th IEEE International Conference on Image Processing (ICIP) . Piscataway, NJ, USA : IEEE , 2018 : 1797 - 1801 .
LI Yue , YI Yan , LIU Dong , et al . Neural-network-based cross-channel intra prediction [J ] . ACM Transactions on Multimedia Computing Communications and Applications , 2021 , 17 ( 3 ): 77 .
MEYER M , WIESNER J , SCHNEIDER J , et al . Convolutional neural networks for video intra prediction using cross-component adaptation [C ] // 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Piscataway, NJ, USA : IEEE , 2019 : 1607 - 1611 .
ZHU Linwei , ZHANG Yun , WANG Shiqi , et al . Deep learning-based chroma prediction for intra versatile video coding [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2021 , 31 ( 8 ): 3168 - 3181 .
BLANCH M G , BLASI S , SMEATON A , et al . Chroma intra prediction with attention-based CNN architectures [C ] // 2020 IEEE International Conference on Image Processing (ICIP) . Piscataway, NJ, USA : IEEE , 2020 : 783 - 787 .
BLANCH M G , BLASI S , SMEATON A F , et al . Attention-based neural networks for chroma intra prediction in video coding [J ] . IEEE Journal of Selected Topics in Signal Processing , 2021 , 15 ( 2 ): 366 - 377 .
ZOU Chengyi , WAN Shuai , JI Tiannan , et al . Spatial information refinement for chroma intra prediction in video coding [C ] // 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . Piscataway, NJ, USA : IEEE , 2021 : 1422 - 1427 .
ZOU Chengyi , WAN Shuai , MRAK M , et al . Towards lightweight neural network-based chroma intra prediction for video coding [C ] // 2022 IEEE International Conference on Image Processing (ICIP) . Piscataway, NJ, USA : IEEE , 2022 : 1006 - 1010 .
ZOU Chengyi , WAN Shuai , JI Tiannan , et al . Chroma intra prediction with lightweight attention-based neural networks [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2024 , 34 ( 1 ): 549 - 560 .
KHAIRAT A , NGUYEN T , SIEKMANN M , et al . Adaptive cross-component prediction for 4∶4∶4 high efficiency video coding [C ] // 2014 IEEE International Conference on Image Processing (ICIP) . Piscataway, NJ, USA : IEEE , 2014 : 3734 - 3738 .
ZHANG Xingyu , GISQUET C , FRANÇOIS E , et al . Chroma intra prediction based on inter-channel correlation for HEVC [J ] . IEEE Transactions on Image Processing , 2014 , 23 ( 1 ): 274 - 286 .
ZHANG Tao , FAN Xiaopeng , ZHAO Debin , et al . Improving chroma intra prediction for HEVC [C ] // 2016 IEEE International Conference on Multimedia & Expo Workshops (ICMEW) . Piscataway, NJ, USA : IEEE , 2016 : 1 - 6 .
LEE S H , MOON J W , BYUN J W , et al . A new intra prediction method using channel correlations for the H.264/AVC intra coding [C ] // 2009 Picture Coding Symposium . Piscataway, NJ, USA : IEEE , 2009 : 1 - 4 .
YEO C , TAN Y , LI Zhengguo , et al . Chroma intra prediction using template matching with reconstructed luma components [C ] // 2011 18th IEEE International Conference on Image Processing . Piscataway, NJ, USA : IEEE , 2011 : 1637 - 1640 .
KIM W S , PU Wei , KHAIRAT A , et al . Cross-component prediction in HEVC [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2020 , 30 ( 6 ): 1699 - 1708 .
ZHANG Kai , CHEN Jianle , ZHANG Li , et al . Multi-model based cross-component linear model chroma intra-prediction for video coding [C ] // 2017 IEEE Visual Communications and Image Processing (VCIP) . Piscataway, NJ, USA : IEEE , 2017 : 1 - 4 .
ZHANG Kai , CHEN Jianle , ZHANG Li , et al . Enhanced cross-component linear model for chroma intra-prediction in video coding [J ] . IEEE Transactions on Image Processing , 2018 , 27 ( 8 ): 3983 - 3997 .
KUO Chewei , LI Xinwei , XIU Xiaoyu , et al . Gradient linear model for chroma intra prediction [C ] // 2023 Data Compression Conference (DCC) . Piscataway, NJ, USA : IEEE , 2023 : 13 - 21 .
ASTOLA P . AHG12: convolutional cross-component model (CCCM) for intra prediction: JVET-Z0064 [EB/OL ] . ( 2022-04-13 ) [ 2024-05-01 ] . https://jvet-experts.org/ https://jvet-experts.org/ .
BROSS B , WANG Yekui , YE Yan , et al . Overview of the versatile video coding (VVC) standard and its applications [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2021 , 31 ( 10 ): 3736 - 3764 .
YU Jiahui , YANG Linjie , XU Ning , et al . Slimmable neural networks [C ] // 7th International Conference on Learning Representations . New York, USA : ICLR , 2019 : 1 - 12 .
YANG Fei , HERRANZ L , CHENG Yongmei , et al . Slimmable compressive autoencoders for practical neural image compression [C ] // 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2021 : 4996 - 5005 .
LIU Zhaocheng , HERRANZ L , YANG Fei , et al . Slimmable video codec [C ] // 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) . Piscataway, NJ, USA : IEEE , 2022 : 1742 - 1746 .
MA Di , ZHANG Fan , BULL D R . BVI-DVC: a training database for deep video compression [J ] . IEEE Transactions on Multimedia , 2022 , 24 : 3847 - 3858 .
TIMOFTE R , AGUSTSSON E , GOOL L V , et al . NTIRE 2017 challenge on single image super-resolution: methods and results [C ] // 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) . Piscataway, NJ, USA : IEEE , 2017 : 1110 - 1121 .
BROWNE A , YE Y , KIM S . Algorithm description for versatile video coding and test model 18(vtm 18): JVET-AB2002 [EB/OL ] . ( 2022-10-28 ) [ 2024-05-01 ] . https://jvet-experts.org/ https://jvet-experts.org/ .
BOYCE J , SUEHRING K , LI L , et al . JVET common test conditions and software reference configurations: JVET-J1010 [EB/OL ] . ( 2018-04-20 ) [ 2024-05-01 ] . https://jvet-experts.org/ https://jvet-experts.org/ .
0
Views
5
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621