Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Audio & Voice

Gemini 3.1 Flash TTS generates expressive multi-speaker speech with natural-language control over tone, pacing, and emotion.
Best for: Creating multilingual narration, Producing multi-speaker dialogue, Testing expressive voice directions
Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Turn a video, social link, or audio file into a transcript, subtitles, and a summary without signing up.
Audio Transcriber AI turns uploaded recordings into searchable text for meetings, podcasts, interviews, and voice notes.