Conceptual

Vision-Language Aligned Diffusion for Text-to-Image Generation

A text-to-image generation framework that improves alignment between complex text prompts and generated images by decomposing prompts into global and local semantic representations and contrastively aligning text and image embeddings from a pretrained vision-language model, then driving a hierarchical multi-stage diffusion model with these aligned representations. Low-rank adaptation keeps the fine-tuning of the large vision-language model computationally efficient.