Advances and Challenges in Deep Learning Based Image Generation Models

Authors

  • Ruizhi Liu School of Fuzhou University, Fuzhou,350108, China

DOI:

https://doi.org/10.54097/h7kbhc75

Keywords:

Computer Vision, Image Generation, Deep Learning.

Abstract

A hot topic in current computer vision research is image generation, which essentially refers to the use of computer algorithms and models to generate artistic and creative images. With the development of artificial intelligence and deep learning, numerous image generation models have emerged, and image generation technology is gradually becoming an important research direction in the fields of computer graphics and art. This has driven the widespread application of image generation in areas such as film, gaming, and painting. This paper provides a preliminary introduction to the basic principles of image generation models, including popular models, their evolution and variants. It then discusses some of the application areas in which image generation can be used. Finally, by analysing the challenges and issues faced by image generation models, this paper looks ahead to some of the work that future researchers will undertake and the direction in which image generation models will develop.

Downloads

Download data is not yet available.

References

[1] Kingma D P, Welling M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.

[2] Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial networks. Communications of the ACM, 2020, 63(11): 139–144.

[3] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020, 33: 6840–6851.

[4] Sohn K, Lee H, Yan X. Learning structured output representation using deep conditional generative models. Advances in Neural Information Processing Systems, 2015, 28: 3483–3491.

[5] Higgins I, Matthey L, Pal A, et al. β-VAE: Learning basic visual concepts with a constrained variational framework. International Conference on Learning Representations, 2017.

[6] Larsen A B L, Sønderby S K, Larochelle H, et al. Autoencoding beyond pixels using a learned similarity metric. Proceedings of the International Conference on Machine Learning, 2016, 48: 1558–1566.

[7] Rombach R, Blattmann A, Lorenz D, et al. High-Resolution Image Synthesis with Latent Diffusion Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 10684–10695.

[8] Radford A, Metz L, Chintala S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.

[9] Gulrajani I, Ahmed F, Arjovsky M, et al. Improved training of Wasserstein GANs. Advances in Neural Information Processing Systems, 2017, 30: 5767–5777.

[10] Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019: 4401–4410.

[11] Ramesh A, Pavlov M, Goh G, et al. Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning, 2021, 139: 8821–8831.

Downloads

Published

30-12-2025

How to Cite

Liu, R. (2025). Advances and Challenges in Deep Learning Based Image Generation Models. Highlights in Science, Engineering and Technology, 160, 60-65. https://doi.org/10.54097/h7kbhc75