I originally intended to include this section in another article1, but found that this small part contains too much content and is somewhat less relevant to the main text theme, so I decided to extract it into a separate article, which also served as a research dive into diffusion models. Here, I will simply introduce the principles and history of Diffusion models, along with my own integration of related knowledge.

What is a Diffusion Model?

For now, let me take a shortcut; for detailed content, you can refer to Professor Li Mu’s video2, or the 3 by Su Shen regarding diffusion models, which covers very deep mathematical principles. I will not elaborate further here.

But in essence, it generates a complete image from random noise or an initial value. During training, the process is reversed: starting from complete images and gradually transforming them into random noise. The first ‘D’ in DDPM stands for Denoising.

A Brief History of Diffusion Models

A Brief History of Generative Models Pushed to the Limit

A Brief History of the Arms Race

  • 2020.6.19 Ho et al. published the DDPM paper titled 《Diffusion Probabilistic Models for Image Generation》4.
  • 2022.4.13 OpenAI published the Dall·E 2 paper titled 《Hierarchical Text-Conditional Image Generation with CLIP Latents》5.
  • 2022.4.13 Stable Diffusion published the paper 《High-Resolution Image Synthesis with Latent Diffusion Models》6.
  • 2022.5.30 Google Brain released Imagen7 8.
  • 2022.6.19 Google & NVIDIA presented the Tutorial 《Denoising Diffusion-based Generative Modeling: Foundations and Applications》 at CVPR 20229.
  • 2022.8.30 OpenAI published a blog post about image inpainting using Dall·E10.

Trial Runs of Various Models

Since I am not very familiar with metrics in the field of image generation, and given that different models offer varying parameters, there are differences in generation performance and iteration counts. I have only conducted some trial runs of these models, so this cannot serve as a strict performance metric; it merely illustrates the demo effects provided by each.

All images were generated using the same Prompt: On a black starry background, Pikachu stands with a star stick in his right hand

The trial addresses for each model are listed below, and these are all free versions:

Generation Results:

Generated by Dall·E-mini

Generated by Dall·E-mini

Generated by ERNIE-ViLG

Generated by ERNIE-ViLG

Generated by Stable Diffusion

Generated by Stable Diffusion

Generated by Stability AI

Generated by Stability AI

In these samples, Baidu’s ERNIE-ViLG on Hugging Face produces high-quality images, but offers no way to adjust its parameters, since Baidu’s ‘Infinite Exploration’ service is not publicly available outside Hugging Face. Stability AI, the company behind Stable Diffusion, also produces fairly clear images, although my parameter choices may have been poor, making its results look slightly worse than ERNIE-ViLG’s. Dall·E-mini and Dream Studio produced less satisfactory results, but broadly captured the intended meaning.

Roughly speaking, besides the pre-trained CLIP model, the parameters that have an impact are the number of iterations and the image size. I don’t know about other parameters, nor am I very good at tuning them –.

Reference