Comparison of Improving Methods for Image Generation Models Targeted on Large-Scale Image Datasets: An Exploration of Improvement Directions for Derived Models Based on Gans, Autoregressive Models, and Diffusion Models

Authors

  • Chenwei Tian School of CHIPS, XJTLU University, Suzhou,215400, China

DOI:

https://doi.org/10.54097/x92mcf49

Keywords:

Image generation models; Generative Adversarial Networks; Diffusion model; Autoregressive; ImageNet.

Abstract

With the rapid progress in the field of deep learning, image generation technology (via GANs, AR, and Diffusion Models) has advanced in computer vision, demonstrating strong synthesis capabilities on large datasets, such as ImageNet (512×512). Yet, people need improvement in generation quality, training efficiency and generalisation. This paper focuses on their derived models, proposing an "improvement object"-based framework at the dataset, training method, and network architecture levels. It identifies key optimisations (architecture, training strategies, dataset expansion) that significantly boost performance. For instance, at the dataset level, the framework integrates multi-source data augmentation. Regarding training methods, it introduces adaptive learning rate scheduling and cross-modal supervision. As for network architecture, a lightweight attention mechanism is embedded to enhance detail capture, resulting in generated 512×512 images achieving 12% higher FID scores than baseline models. Experiments on ImageNet and COCO verify the framework’s universality, providing a systematic solution for advancing image generation. The article aims to provide systematic references and guide the development of generative AI in high-resolution, large-scale scenarios.

Downloads

Download data is not yet available.

References

[1] Ren S, Yu Q, He J, Shen X, Yuille A, Chen L-C. Beyond next-token: next-X prediction for autoregressive visual generation. arXiv:2502.20388, 2025.

[2] Alonso E, Moysset B, Messina R. Adversarial generation of handwritten text images conditioned on sequences. Proceedings of the International Conference on Document Analysis and Recognition, 2019: 481-486.

[3] Fogel S, Averbuch-Elor H, Cohen S, Mazor S, Litman R. ScrabbleGAN: semi-supervised varying length handwritten text generation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020: 4324–4333.

[4] Zdenek J, Nakayama H. JokerGAN: memory-efficient model for handwritten text generation with text line awareness. Proceedings of the 29th ACM International Conference on Multimedia, 2021: 5655–5663.

[5] Liu X, Meng G, Xiang S, Pan C. Handwritten text generation via disentangled representations. IEEE Signal Processing Letters, 2021, 28: 1838–1842.

[6] Kang L, Riba P, Wang Y, Rusinol M, Fornes A, Villegas M. GANwriting: content-conditioned generation of styled handwritten word images. European Conference on Computer Vision, 2020, 12368: 237–289.

[7] Gan J, Wang W. HiGAN: handwriting imitation conditioned on arbitrary-length texts and disentangled styles. AAAI Conference on Artificial Intelligence, 2021, 35(9): 7484–7492.

[8] Davis B, Tensmeyer C, Price B, Wigington C, Morse B, Jain R. Text and style conditioned GAN for generation of offline handwriting lines. arXiv:2009.00678, 2020.

[9] Gan J, Wang W, Leng J, Gao X. HiGAN+: handwriting imitation GAN with disentangled representations. ACM Transactions on Graphics, 2022, 42(1): 1–17.

[10] Luo C, Zhu Y, Jin L, Li Z, Peng D. SLOGAN: handwriting style synthesis for arbitrary-length and out-of-vocabulary text. IEEE Transactions on Neural Networks and Learning Systems, 2022.

[11] Ganz R, Elad M. BIGRoC: boosting image generation via a robust classifier. Transactions on Machine Learning Research, 2023.

[12] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020, 33: 6840–6851.

[13] Dhariwal P, Nichol A. Diffusion models beat GANs on image synthesis. arXiv:2105.05233, 2021.

[14] Brock A, Donahue J, Simonyan K. Large scale GAN training for high fidelity natural image synthesis. arXiv:1809.11096.

[15] Gao S, Zhou P, Cheng M-M, Yan S. MDTv2: masked diffusion transformer is a strong image synthesizer. arXiv:2303.14389, 2023.

[16] Sauer A, Schwarz K, Geiger A. StyleGAN-XL: scaling StyleGAN to large diverse datasets. ACM SIGGRAPH Conference Proceedings, 2022.

[17] Sun P, Jiang Y, Lin T. Unified continuous generative models. arXiv:2505.07447, 2025.

[18] Miyato T, Kataoka T, Koyama M, Yoshida Y. Spectral normalization for generative adversarial networks. International Conference on Learning Representations (ICLR), 2018.

[19] Chen T, Zhai X, Ritter M, Lucic M, Houlsby N. Self-supervised GANs via auxiliary rotation loss. Proceedings of the Conference on Computer Vision and Pattern Recognition, 2019.

[20] Lee K S, Tran N T, Cheung N M. InfoMax-GAN: improved adversarial image generation via information maximization and contrastive learning. Winter Conference on Applications of Computer Vision, 2021.

[21] Esser P, Rombach R, Ommer B. Taming transformers for high-resolution image synthesis. Conference on Computer Vision and Pattern Recognition, 2021.

[22] Chang H, Zhang H, Jiang L, Liu C, Freeman W T. Masked generative image transformer. Conference on Computer Vision and Pattern Recognition, 2022.

[23] Peebles W, Xie S. Scalable diffusion models with transformers. International Conference on Computer Vision, 2023.

[24] Tian K, Jiang Y, Yuan Z, Peng B, Wang L. Visual autoregressive modeling: scalable image generation via next-scale prediction. Advances in Neural Information Processing Systems, 2024.

[25] Yu S, Kwak S, Jang H, Jeong J, Huang J, Shin J, Xie S. Representation alignment for generation: training diffusion transformers is easier than you think. arXiv:2410.06940, 2024.

[26] Ma N, Goldstein M, Albergo M S, Boffi N M, Vanden-Eijnden E, Xie S. SIT: exploring flow and diffusion-based generative models with scalable interpolant transformers. European Conference on Computer Vision, 2024: 23–40.

[27] Zheng H, Nie W, Vahdat A, Anandkumar A. Fast training of diffusion models with masked transformers. arXiv:2306.09305, 2023.

[28] Karras T, Aittala M, Lehtinen J, Hellsten J, Aila T, Laine S. Analyzing and improving the training dynamics of diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024: 24174–24184.

[29] Wang S, Tian Z, Huang W, Wang L. DDT: decoupled diffusion transformer. arXiv:2504.05741, 2025.

[30] Yao J, Yang B, Wang X. Reconstruction vs. generation: taming optimization dilemma in latent diffusion models. arXiv:2501.01423, 2025.

[31] Chen T, Zhai X, Ritter M, et al. Self-Supervised GANs via Auxiliary Rotation Loss. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2019.

[32] Atito S, Awais M, Kittler J. SIT: Self-supervised vision transformer. arXiv preprint arXiv:2104.03602, 2021.

[33] Kang M, Zhu J-Y, Zhang R, et al. Scaling up GANs for text-to-image synthesis. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.

[34] Kingma D P, Ba J. Adam: A method for stochastic optimization. In: International Conference on Learning Representations (ICLR), 2015.

[35] Kingma D P, Welling M. Auto-encoding variational bayes. In: International Conference on Learning Representations (ICLR), 2014.

[36] Kynkäänniemi T, Aittala M, Karras T, et al. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. arXiv preprint arXiv:2404.07724, 2024.

[37] Lee D, Kim C, Kim S, et al. Autoregressive image generation using residual quantization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.

[38] Lipman Y, Chen R T Q, Ben-Hamu H, et al. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.

[39] Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 2011, 12: 2825–2830.

Downloads

Published

30-12-2025

How to Cite

Tian, C. (2025). Comparison of Improving Methods for Image Generation Models Targeted on Large-Scale Image Datasets: An Exploration of Improvement Directions for Derived Models Based on Gans, Autoregressive Models, and Diffusion Models. Highlights in Science, Engineering and Technology, 160, 88-99. https://doi.org/10.54097/x92mcf49