2501.00794
VoiceRestore is a single self-supervised model that restores degraded speech recordings, removing background noise, reverberation, compression artifacts, and bandwidth limitations within one unified …
A single unified model restores degraded speech -- background noise, reverberation, compression artifacts, and bandwidth limitation -- by learning a conditional flow-matching vector field with a Transformer that transports a degraded speech spectrogram toward a clean one, steered at inference by classifier-free guidance. Learners see how a continuous-time generative transport replaces paired supervised training: the model is trained only on synthetic degradations of clean speech, needs no paired clean/degraded corpus, and generalizes to real short utterances and to long-form monologues and dialogues.
VoiceRestore is a single self-supervised model that restores degraded speech recordings, removing background noise, reverberation, compression artifacts, and bandwidth limitations within one unified …