Home /permanent

Audio Compression Model

An Audio Compression Model is a model that compresses an audio waveform into a compact representation and can decode that representation back into audio.

Modern neural versions, like SoundStream and Encodec, use an encoder, a quantiser (typically Residual Vector Quantisation) and a decoder, trained end-to-end. See Neural Audio Codec.

As well as reducing bitrate, the discrete codes they produce can be used as tokens for audio language models, like in AudioLM: a Language Modeling Approach to Audio Generation.