Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Audio & Voice

Qwen3-TTS is an open multilingual speech family for voice cloning, expressive generation, streaming, and instruction-based delivery.
Best for: Cloning a reference voice, Generating multilingual speech, Building low-latency voice applications
Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Turn a video, social link, or audio file into a transcript, subtitles, and a summary without signing up.
Audio Transcriber AI turns uploaded recordings into searchable text for meetings, podcasts, interviews, and voice notes.