MusicGen
Meta
Meta's AI music generation model available in 300M, 1.5B, and 3.3B parameter sizes. Single-stage auto-regressive Transformer trained over 32kHz EnCodec tokenizer with 4 codebooks sampled at 50Hz. Trained on 20K hours of licensed music from internal dataset (10K high-quality tracks) plus ShutterStock and Pond5 collections (390K instrument-only tracks). Eliminates cascading model requirement through efficient token interleaving - only 50 auto-regressive steps per second of audio. Supports both text-to-music and melody-guided generation. Evaluated on MusicCaps benchmark showing superiority vs baselines. Released April-May 2023. Part of AudioCraft toolkit alongside AudioGen and EnCodec.
Strengths
- Text-to-music and melody-guided generation modes
- Trained on 20K hours of high-quality licensed music
- Single-stage architecture eliminates cascading complexity
- Efficient generation - only 50 auto-regressive steps per second
- Available in three sizes (300M, 1.5B, 3.3B) for different use cases
- MIT license enables full commercial use
- Part of comprehensive AudioCraft toolkit
Caveats
- Released June 2023 - newer music models may surpass quality
- Training limited to 20K hours vs some competitors
- Primarily instrument-focused - vocals less refined
- Generated music may lack coherence for longer compositions
- Consider Suno, Udio, or other specialized music models for production
- Vendor
- Meta
- License
- MIT
- Release Date
- 2023-06-01
- Modalities
- audio
Capabilities
Resources
Pricing
Open source (MIT) - free. Requires GPU. Available on Hugging Face and Replicate