Home /permanent

Acoustic Tokens

Acoustic Tokens are discrete tokens produced by a neural audio codec, such as SoundStream or Encodec, that capture the fine acoustic details of an audio waveform: speaker identity, recording conditions and so on.

Because the codec is trained to reconstruct audio, acoustic tokens allow high-quality synthesis. However, a language model trained only on acoustic tokens struggles with long-term structure (for speech, it tends to produce babbling).

AudioLM: a Language Modeling Approach to Audio Generation combines them with semantic tokens, which capture long-term structure, to get both. See Audio Tokenisation and Residual Vector Quantisation.