Conceptual

Sound-Guided Visual Scene Editing with Latent Diffusion Models

SoundBrush edits existing images and 3D scenes using sound as the control signal, extending a text-to-image latent diffusion model by learning a mapping network that projects audio features into the model's textual token space. Because paired sound-and-edited-image data is scarce, it is trained on a synthetically constructed dataset assembled from off-the-shelf sound-source localization, image inpainting, and prompt-based editing models, and can insert sounding objects or adjust scene appearance to match diverse in-the-wild audio while preserving the original content.