Home /permanent

VALL-E

VALL-E is a zero-shot TTS language model based on RVQ. See Neural Codec Language Models are Zero-Shot Text-to-Speech Synthesizers VALL-E converts text to phonemes and a three-second speech prompt to codec tokens. A neural codec language model predicts tokens that an audio decoder turns into personalized speech.