Home /permanent

Neural Audio Codec

Neural Audio Codec is an audio codec that uses neural networks to compress and decompress audio, instead of the hand-designed signal processing used by traditional codecs like Opus Audio Codec and EVS.

A typical design has three parts: an encoder that turns the waveform into a compact sequence of embeddings, a quantiser (often Residual Vector Quantisation) that turns the embeddings into discrete codes, and a decoder that reconstructs the waveform from the codes. They're usually trained with a mix of reconstruction and adversarial losses.

Examples include Soundstream: An End-to-End Neural Audio Codec, Encodec and High-Fidelity Audio Compression with Improved RVQGAN.

Since they turn audio into discrete tokens, they're also used as audio tokenisers for audio language models like AudioLM: a Language Modeling Approach to Audio Generation and MusicGen.