Conceptual

Adapter-Tuned Self-Supervised Features for Zero-Shot Voice Conversion

A zero-shot voice conversion method that disentangles linguistic content from speaker style by attaching trainable adapters to a frozen self-supervised speech model, learning weighted combinations of its intermediate-layer features under auxiliary objectives. The disentangled content and speaker representations are fused by a conditional-flow-matching decoder with optimal-transport paths and cross-attention speaker conditioning to synthesize high-quality converted speech.