A Prediction Model for Video Frames of Road Conditions Using Residual Blocks Generating Adversarial Networks[J]. 2018, 52(10): 146-152+166.
DOI:
A Prediction Model for Video Frames of Road Conditions Using Residual Blocks Generating Adversarial Networks[J]. 2018, 52(10): 146-152+166.DOI: 10.7652/xjtuxb201810020.
A Prediction Model for Video Frames of Road Conditions Using Residual Blocks Generating Adversarial Networks
A novel model to predict video frames of road conditions is proposed to solve the problem that in the prediction area of road condition video frames
most of predicting models are of such as low resolution
image blurring and lack of local details
and the model uses residual blocks to generate adversarial networks. Multi-cascade residual units are used to extract feature maps preliminarily and two discriminators are adopted to raise the quality of images generated by the GANs. The model power to extract feature maps is improved by using perceptual loss network
and the Adam optimizer is used to update network weights. The proposed model produces future frames in the next moment which have great temporal coherency with the input frames through using semi-supervised deep learning framework. The model is trained and tested using datasets KITTI
and the results show that the generated frames by the proposed model have the resolution of 256×512
which are 2 to 4 times higher than the results from pixel mean based functions and past foreign works
the sharpness accuracy increases by 1 to 2 magnitudes
and the generated frames are more responsive to human visual perception. The frame images predicted by the model are of higher quality and more practical value
and provide more effective features for other downstream models such as detection algorithms.
关键词
Keywords
references
SRIVASTAVA N, MANSIMOV E, SALAKHUDINOV R. Unsupervised learning of video representations using LSTMS [C]∥Proceedings of the 32nd International Conference on Machine Learning. New York, USA: ACM, 2015: 843-852.
HOCHREITER S, SCHMIDHUBER J. Long short-term memory [J]. Neural Computation, 1997, 9(8): 1735-1780.
BHATTACHARJEE P, DAS S, BHATTACHARJEE P, et al. Temporal coherency based criteria for predicting video frames using deep multi-stage generative adversarial networks [C]∥Proceedings of the Advances in Neural Information Processing Systems. New York, USA: NIPS, 2017: 4271-4280.
MATHIEU M, COUPRIE C, LECUN Y. Deep multi-scale video prediction beyond mean square error [EB/OL].(2015-11-17)[2018-02-28]. https:∥arxiv.org/abs/1511.05440.
LIANG X, LEE L, DAI W, et al. Dual motion GAN for future-flow embedded video prediction [C]∥Proceedings of the IEEE International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2017: 1762-1770.
XIONG W, LUO W, MA L, et al. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks [EB/OL].(2017-09-22)[2018-02-28]. https:∥arxiv.org/abs/1709.07592.
LEDIG C, THEIS L, HUSZÁR F, et al. Photo-realistic single image super-resolution using a generative adversarial network [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 105-114.
CHEN Q, KOLTUN V. Photographic image synthesis with cascaded refinement networks [C]∥Proceedings of the IEEE International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2017: 1520-1529.
YANG C, LU X, LIN Z, et al. High-resolution image inpainting using multi-scale neural patch synthesis [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 4076-4084.
GOODFELLOW I J, POUGET-ABADIE J, MIRZA M, et al. Generative adversarial networks [C]∥Proceedings of the Advances in Neural Information Processing Systems. New York, USA: NIPS, 2014: 2672-2680.
MIRZA M, OSINDERO S. Conditional generative adversarial nets [EB/OL].(2014-11-06)[2018-02-28]. https:∥arxiv.org/abs/1411.1784.
ISOLA P, ZHU J Y, ZHOU T, et al. Image-to-image translation with conditional adversarial networks [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5967-5976.
RONNEBERGER O, FISCHER P, BROX T. U-net: convolutional networks for biomedical image segmentation [C]∥Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Berlin, Germany: Springer, 2015: 234-241.
HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognitio. Piscataway, NJ, USA: IEEE, 2016: 770-778.
WANG T C, LIU M Y, ZHU J Y, et al. High-resolution image synthesis and semantic manipulation with conditional GANs [EB/OL].(2017-11-30)[2018-02-25]. https:∥arxiv.org/abs/1711.11585.
JOHNSON J, ALAHI A, LI F F. Perceptual losses for real-time style transfer and super-resolution [C]∥Proceedings of the European Conference on Computer Vision. Berlin, Germany: Springer, 2016: 694-711.
SIMONYAN K, ZISSERMAN A. Very deep convolutional networks for large-scale image recognition [EB/OL].(2014-09-04)[2018-02-24]. https:∥arxiv.org/abs/1409.1556.
GEIGER A, LENZ P, STILLER C, et al. Vision meets robotics: the KITTI dataset [J]. Journal of Robotics Research, 2013, 32(11): 1231-1237.
ULYANOV D, VEDALDI A, LEMPITSKY V, et al. Instance normalization: the missing ingredient for fast stylization [EB/OL].(2016-07-27)[2018-03-10]. https:∥arxiv.org/abs/1607.08022.
ARJOVSKY M, CHINTALA S, BOTTOU L. Wasserstein generative adversarial networks [C]∥Proceedings of the International Conference on Machine Learning. New York, USA: ACM, 2017: 214-223.
KINGMA D P, BA J. Adam: a method for stochastic optimization [EB/OL].(2014-12-22)[2018-02-20]. https:∥arxiv.org/abs/1412.6980.