From Gans to Diffusion Models: Text-To-Image Generation
DOI:
https://doi.org/10.54097/79d59267Keywords:
GANs; Diffusion Models; text-to-image generation.Abstract
This paper traces the evolution of text-to-image generation (TIG) techniques from Generative Adversarial Networks (GANs) to Diffusion Models (DMs). It first introduces GAN variants, including DCGAN, WGAN, MGGAN, and StyleGAN. While these popular GANs pioneered image synthesis through adversarial training of the generator and discriminator, they suffered from training instability, mode collapse, and the lack of diversity. Therefore, the systematic introduction of representative DMs—including DDPM, Guided Diffusion, GLIDE, Stable Diffusion, and Imagen—shows how they address these issues through iteratively denoising, achieving unprecedented image fidelity, semantic alignment, and generation stability. Quantitative comparisons on datasets such as COCO and CUB show that DMs consistently outperform GANs in metrics like FID, IS, and CLIP score, though GANs retain shorter inference time. Nevertheless, critical challenges such as generation efficiency, understanding of complex prompts, and safety controls remain. This paper analyses possible reasons for those problems while pointing out key directions for future work.
Downloads
References
[1] Iqbal T, Qureshi S. The survey: Text generation models in deep learning. Journal of King Saud University-Computer and Information Sciences, 2022, 34(6): 2515–2528.
[2] Bang D, Shim H. MGGAN: Solving mode collapse using manifold guided training. arXiv preprint arXiv:1804.04391, 2018.
[3] Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv preprint arXiv:1503.03585, 2015.
[4] Zhou J. Text-to-image generation: A literature review focusing on the diffusion model. ITM Web of Conferences, 2025, 73: 02037.
[5] Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial nets. Advances in Neural Information Processing Systems, 2014, 27.
[6] Arjovsky M, Chintala S, Bottou L. Wasserstein GAN. arXiv preprint arXiv:1701.07875, 2017.
[7] Gulrajani I, Ahmed F, Arjovsky M, et al. Improved training of Wasserstein GANs. arXiv preprint arXiv:1704.00028, 2017.
[8] Radford A, Metz L, Chintala S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.
[9] Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019: 4401–4410.
[10] Dieng A B, Ruiz F J R, Blei D M, Titsias M K. Prescribed GAN (PresGAN): Enabling explicit density estimation. arXiv preprint arXiv:1910.04302, 2019.
[11] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. arXiv preprint arXiv:2006.11239, 2020.
[12] Dhariwal P, Nichol A. Diffusion models beat GANs on image synthesis. arXiv preprint arXiv:2105.05233, 2021.
[13] Nichol A, Dhariwal P, Ramesh A, et al. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
[14] Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022: 10684–10695.
[15] Saharia C, Chan W, Saxena S, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 2022.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Highlights in Science, Engineering and Technology

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







