Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Search by what you want to achieve. Strong free badges are reserved for tools whose access was actually checked.
75 tools found
Stable Audio 3.0 provides open audio-generation models for creating, extending, editing, and inpainting music and sound.
Turn a video, social link, or audio file into a transcript, subtitles, and a summary without signing up.
Audio Transcriber AI turns uploaded recordings into searchable text for meetings, podcasts, interviews, and voice notes.
Audio Cut trims common audio formats in the browser for quick cleanup, ringtones, and shorter recordings.
Qwen3-TTS is an open multilingual speech family for voice cloning, expressive generation, streaming, and instruction-based delivery.
Fineshare Singify creates complete songs, instrumentals, covers, and cloned singing voices from prompts, lyrics, or reference media.
High-quality neural text-to-speech synthesis supporting dozens of global languages for cost-effective audio production.
It helps you remake a voice from 5 seconds of audio and produce real-time speech synthesis with Chatterbox Turbo (an open source model).
An open-source, full-stack lyric-to-audio engine capable of synthesizing complete five-minute tracks with synchronized vocal and instrumental perform...
It helps you clone any voice in seconds and control the emotion of the speech maked.
Microsoft's 1.5 billion parameter transformer model specialized for high-fidelity, multi-speaker conversational speech synthesis.
It helps you make high-quality music and sounds up to 3 minutes long in less than 2 seconds.
It helps you elevenLabs MCP is an open source Python package that makes it easy to integrate ElevenLabs voice capabilities into your applications.
A high-performance open-source text-to-speech model providing granular control over vocal characteristics and custom cloning.
A localized watermarking and detection system for pinpointing synthetic audio segments within complex audio streams in real-time.
Lightning-fast, browser-based transcription for over 100 languages with precision-timed subtitle exports.
It turns your text into high-quality speech and lets you customise voice profiles, speed and pitch.