SongGeneration is an open-source music foundation model that aims to revolutionize AI music generation. We develop a high-quality, controllable, and commercially viable system for creating original music. LeVo 2 leverages a novel hybrid LLM-Diffusion architecture to balance musicality and audio quality, achieving state-of-the-art results in perceptual metrics. The core problem it solves is the limitation of previous models in generating music with commercial-grade quality and accurate lyrics.
SongGeneration 2 achieves commercial-grade musicality, significantly outperforming open-source baselines. It demonstrates exceptional controllability through multi-modal instructions and supports full-length song generation with vocals and accompaniment. A key architectural innovation is the combination of an LLM and diffusion model for superior results.
- Multi-Track Generation: Supports generation of pure music, pure vocals, and dual-track (vocals + accompaniment) outputs.
- High Controllability: Responds to multi-modal instructions (text, audio) for precise control.
- Low Memory Usage: Model can run with as little as 10GB of GPU memory.
- Advanced Architecture: Combines LLM and Diffusion models for improved musicality and sound quality.
- Improved Performance: Achieves high scores in Parallel Evaluation Framework (PER) and other metrics.
- Flexible Output: Supports generation of pure music, vocals, and dual-track outputs.
- Enhanced Training: Utilizes automated aesthetics evaluation and multi-stage post-training for high quality.
SongGeneration 2 has recently been released with several model versions available on Hugging Face Hub. The project has a growing community and active development, with ongoing efforts to improve performance and add new features. The release of core components like the data processing pipeline and finetuning scripts indicates a maturing ecosystem. Several TODOs have been addressed, including low memory usage models and full-time models.
SongGeneration is beneficial for musicians, content creators, and researchers seeking to generate high-quality music. It addresses real-world use cases such as background music creation, songwriting assistance, and experimental music production. The project provides a powerful alternative to manual music creation and other less sophisticated AI music tools by offering a combination of control, quality, and accessibility.
