Deep Generative Models for Image Synthesis: GANs, VAEs, and Diffusion Approaches

Authors

  • Zhuolin Li Nanjing Foreign Language School FanShan Campus, Nanjing, 210000, China
  • Xiangyu Liu Xiamen JiuXi Senior High School, Xiamen, 361000, China
  • Yibo Lu Nanjing University of Posts and Telecommunications, Nanjing, 210000, China

DOI:

https://doi.org/10.54097/0v5rz196

Keywords:

Image Synthesis; Image Super-Resolution; Image Enhancement and Restoration.

Abstract

This paper provides a comprehensive overview of the evolution and current state of deep generative models in image synthesis, with a focus on Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and diffusion models. It identifies persistent challenges, including difficulty in generating high-resolution details, training instabilities, and limited controllability. The analysis demonstrates that Diffusion Models have emerged as the dominant paradigm, effectively overcoming previous limitations through their robust training dynamics and superior detail preservation. By modelling data distribution via gradual noise addition and removal, diffusion models achieve state-of-the-art image quality while maintaining structural coherence. Their integration with text-conditioning mechanisms enables precise semantic control, facilitating diverse applications such as high-fidelity text-to-image synthesis, medical image enhancement, and accelerating creative design. The review further examines practical implementations, including SRGAN for super-resolution, CVAE for image restoration, and diffusion-based text-to-image systems, highlighting their real-world impact across healthcare, media, and scientific imaging. The work concludes that diffusion models, particularly with latent-space optimisation, represent the most promising direction for future research, balancing unprecedented quality with increasing computational efficiency to enable broader deployment in critical domains.

Downloads

Download data is not yet available.

References

[1] Ledig C, Theis L, Huszár F, et al. Photo-realistic single-image super-resolution using a generative adversarial network. IEEE Trans Pattern Anal Mach Intell, 2017, 40(2): 4681–4690.

[2] Dong C, Loy C C, He K, et al. Learning a deep convolutional network for image super-resolution. IEEE Trans Image Process, 2016, 25(2): 614–627.

[3] Yu Y, Ma X, Zheng Y, et al. Deep learning for historical document image super-resolution. Int J Doc Anal Recognit, 2020, 23(4): 395–411.

[4] Baur C, Albarqouni S, Navab N. Super-resolution for medical imaging: A survey. IEEE Trans Med Imaging, 2021, 40(9): 2449–2464.

[5] Liu J, Huang Y, Zhang W, et al. Real-time face super-resolution via generative adversarial networks for surveillance applications. IEEE Trans Circuits Syst Video Technol, 2019, 29(12): 3587–3600.

[6] Wang X, Yu K, Wu S, et al. ESRGAN: Enhanced super-resolution generative adversarial networks. Neurocomputing, 2020, 389: 107–119.

[7] Huang S, Chen Y, Lu W, et al. Remote sensing image super-resolution using generative adversarial networks. ISPRS J Photogramm Remote Sens, 2020, 162: 155–166.

[8] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. J Mach Learn Res, 2021, 22(1): 1–45.

[9] Kingma D P, Welling M. Auto-encoding variational bayes. Found Trends Mach Learn, 2019, 12(4): 307–392.

[10] Saharia C, Chan W, Saxena S, et al. Photorealistic text-to-image diffusion models with deep language understanding. Nat Mach Intell, 2023, 5: 545–556.

[11] Sohn K, Lee H, Yan X. Learning structured output representation using deep conditional generative models. Neural Comput Appl, 2018, 29(5): 1195–1204.

[12] Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models. IEEE Trans Pattern Anal Mach Intell, 2023, 45(1): 10684–10695.

[13] Yu L, Zhao W, Zhang Z, et al. Multimodal learning with diffusion models: A comprehensive survey. Inf Fusion, 2024, 98: 101933.

[14] Iizuka S, Simo-Serra E, Ishikawa H. Globally and locally consistent image completion. ACM Trans Graph, 2017, 36(4): 1–14.

[15] Bao P, Liao J, Song Q, et al. CVAE-GAN: Fine-grained image generation through asymmetric training. Pattern Recognit, 2018, 79: 389–401.

[16] Liu Y, Wang C, Zhang H, et al. Underwater image enhancement via conditional variational autoencoders with multi-scale feature fusion. IEEE Trans Image Process, 2022, 31: 2985–2997.

[17] Wang R, Huang Z, Li Y, et al. Low-light image enhancement via conditional variational autoencoders with illumination-aware constraints. IEEE Trans Image Process, 2020, 29: 6872–6885.

[18] Wang R, Huang Z, Li Y, et al. Medical image restoration via conditional variational autoencoders with attention mechanisms. IEEE Trans Med Imaging, 2021, 40(8): 2176–2187.

[19] Chen X, Liu Y, Feng R, et al. Real-world applications of text-to-image generation: Case studies and lessons learned. ACM Trans Comput Hum Interact, 2023, 30(2): 1–24.

Downloads

Published

30-12-2025

How to Cite

Li, Z., Liu, X., & Lu, Y. (2025). Deep Generative Models for Image Synthesis: GANs, VAEs, and Diffusion Approaches. Highlights in Science, Engineering and Technology, 160, 205-211. https://doi.org/10.54097/0v5rz196