Conceptual

Self-Supervised Speech Restoration with Flow-Matching Transformers

A single unified model restores degraded speech -- background noise, reverberation, compression artifacts, and bandwidth limitation -- by learning a conditional flow-matching vector field with a Transformer that transports a degraded speech spectrogram toward a clean one, steered at inference by classifier-free guidance. Learners see how a continuous-time generative transport replaces paired supervised training: the model is trained only on synthetic degradations of clean speech, needs no paired clean/degraded corpus, and generalizes to real short utterances and to long-form monologues and dialogues.