Simple and Controllable Music Generation
These are my notes from paper Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, Alexandre Défossez.
Introduces MusicGen to tackle conditional music generation: a Language Model that operates over Residual Vector Quantisation tokens.
Comprised of a single-stage transformer LM together with efficient token interleaving patterns: * eliminates the need for cascading several streams * or cascading approaches like Hierarchical Model or Up-sampling. * Includes an algorithm for efficient Token Interleaving Patterns so they don't need additional models for upsampling. * Can generate in mono and stereo. * Conditioned on text descriptions or melodic features.