Advances and Challenges in Deep Learning Based Image Generation Models
DOI:
https://doi.org/10.54097/h7kbhc75Keywords:
Computer Vision, Image Generation, Deep Learning.Abstract
A hot topic in current computer vision research is image generation, which essentially refers to the use of computer algorithms and models to generate artistic and creative images. With the development of artificial intelligence and deep learning, numerous image generation models have emerged, and image generation technology is gradually becoming an important research direction in the fields of computer graphics and art. This has driven the widespread application of image generation in areas such as film, gaming, and painting. This paper provides a preliminary introduction to the basic principles of image generation models, including popular models, their evolution and variants. It then discusses some of the application areas in which image generation can be used. Finally, by analysing the challenges and issues faced by image generation models, this paper looks ahead to some of the work that future researchers will undertake and the direction in which image generation models will develop.
Downloads
References
[1] Kingma D P, Welling M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
[2] Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial networks. Communications of the ACM, 2020, 63(11): 139–144.
[3] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020, 33: 6840–6851.
[4] Sohn K, Lee H, Yan X. Learning structured output representation using deep conditional generative models. Advances in Neural Information Processing Systems, 2015, 28: 3483–3491.
[5] Higgins I, Matthey L, Pal A, et al. β-VAE: Learning basic visual concepts with a constrained variational framework. International Conference on Learning Representations, 2017.
[6] Larsen A B L, Sønderby S K, Larochelle H, et al. Autoencoding beyond pixels using a learned similarity metric. Proceedings of the International Conference on Machine Learning, 2016, 48: 1558–1566.
[7] Rombach R, Blattmann A, Lorenz D, et al. High-Resolution Image Synthesis with Latent Diffusion Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 10684–10695.
[8] Radford A, Metz L, Chintala S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.
[9] Gulrajani I, Ahmed F, Arjovsky M, et al. Improved training of Wasserstein GANs. Advances in Neural Information Processing Systems, 2017, 30: 5767–5777.
[10] Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019: 4401–4410.
[11] Ramesh A, Pavlov M, Goh G, et al. Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning, 2021, 139: 8821–8831.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Highlights in Science, Engineering and Technology

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







