Conceptual
Login

About How to Build Image Generation Models

Copied

Building generative image models from convolution foundations through VAEs and GANs to diffusion and Diffusion Transformers. Covers where convolution survives in a modern latent-diffusion stack and where it was replaced.

Estimated Time to Complete

Only available after login

What You'll Learn

Concepts:
Compute-to-FID Scaling as the Evidence That Ended the Convolutional Denoiser The Convolutional Encoder and Decoder Inside a Latent Diffusion Model Cross-Attention Mechanism Frechet Inception Distance for Image Generation Deterministic ODE Samplers and DDIM Step Skipping Vision Transformer (ViT) Denoising Diffusion Probabilistic Generative Models Assembling a Current Text-to-Image Stack from Text Encoder, VAE, and Transformer Denoiser Distilling a Many-Step Diffusion Sampler into a Few-Step Student Convolutional Neural Networks Diffusion Transformer Architecture in Deep Learning Convolutional Denoisers in Small and On-Device Image Generators The Forward Noising Process That Destroys an Image into Gaussian Noise The Convolution Operation as a Sliding Filter over an Image Cascaded Pixel-Space Super-Resolution Stages Above a Latent Generator Multimodal Diffusion Transformers with Separate Text and Image Streams Adversarial Training as the Default Route to Photorealistic Image Synthesis Self Attention Mechanism in Transformers Receptive Field Growth Through Stacked Convolutional Layers The Latent Space as a Compressed Coordinate System for Images Feature Hierarchies from Edges to Object Parts in Deep Vision Networks Patchifying a Latent Tensor into Transformer Tokens Generative Adversarial Networks in Deep Learning Denoising Diffusion Models in Deep Learning Pooling and Strided Downsampling in Convolutional Networks Classifier-Free Guidance in Conditional Generative Models StyleGAN Style-Based Generation and Latent Disentanglement Autoencoders in Deep Learning Deep Convolutional GAN Generators Built from Transposed Convolutions Noise Schedules in Diffusion Models Adversarial Losses Inside Few-Step Diffusion Distillation Few-Step Sampling and the Function-Evaluation Budget U-Net Encoder-Decoder Architecture for Image Segmentation Rectified Flow Generative Models Digital Images as Pixel Grids and Multi-Channel Tensors Convolutional Inductive Bias as a Requirement for Image Generation Contrastive Language-Image Pretraining (CLIP) Transposed Convolution for Learned Upsampling Mode Collapse in Adversarial Training Many-Step Ancestral Sampling from a Diffusion Model Stochastic Interpolants as a Design Space over Diffusion and Flow Stride and Padding as Controls on Convolution Output Size Learned Convolution Kernels as Edge and Texture Detectors Conditional Flow Matching for Generative Models The U-Net Used as the Denoiser in a Diffusion Model Variational Autoencoders in Deep Learning Diffusion Trained Directly in Pixel Space Latent Diffusion and Stable Diffusion for Text-to-Image Generation

What you will learn

No introduction video available

About Zoe Graystone

Z

Guide profile coming soon.