Conceptual

Vision-Foundation-Model-Aligned Tokenizers for Latent Diffusion

Resolving the reconstruction-generation trade-off in two-stage latent diffusion models by regularizing the visual tokenizer's (VAE) latent space to align with features from a pretrained self-supervised vision foundation model. This tames the difficulty of learning unconstrained high-dimensional latents, expands the reconstruction-generation frontier, and lets high-dimensional Diffusion Transformers converge far faster while reaching state-of-the-art image-generation quality.