Speaker Disentanglement
Speaker Disentanglement is separating the speaker's identity (their timbre and other voice characteristics) from the content of what's being said in a speech representation.
It's the key step in most Voice Conversion approaches: once content and speaker are separated, you can combine the content with a different speaker's characteristics. It's hard to do without losing some of the content, which is the problem ContentVec tries to fix by modifying HuBERT.