Home /permanent

Voice Conversion

Voice Conversion is the task of transforming a recording of one speaker so that it sounds like it was spoken by a different target speaker, while keeping the linguistic content (what was said) the same.

See Voice Conversion.

Most approaches try to separate the content of the speech from the speaker's identity (see Speaker Disentanglement), then combine the content with the target speaker's characteristics and turn it back into audio with a vocoder like HiFi-GAN. It's closely related to Speech Synthesis, except the input is speech rather than text.

I played around with voice conversion in Making Song Covers With My AI Voice.