Conceptual

Test-Time Spatial Constraint Enforcement for Controllable Text-to-Image Generation

A training-free inference-time technique that makes a pretrained text-to-image diffusion model follow complex spatial layouts given with natural-language prompts. It decouples a spatial constraint into a semantic condition and a geometric condition and enforces each separately: prompt completion plus attention-and-word-distance suppression of distracting tokens for semantics, and Region-of-Interest latent relocation with diffusion-based latent-refill for geometry.