An open-source latent diffusion model that synthesizes complete songs with vocals and instrumentation in under ten seconds.
Audio & Voice

It helps you make dubbing, narration and dialogue with natural, expressive rendering.
Best for: podcasters, educators and app developers needing natural voices
An open-source latent diffusion model that synthesizes complete songs with vocals and instrumentation in under ten seconds.
It writes full songs from a text prompt, including lyrics and vocals, with a license for commercial use.
It composes songs on the fly after you pick a mood, genre and tempo, then lets you rearrange them.