It creates realistic custom voices from a few minutes of recorded speech and lets you control emotion and language.
Search by what you want to achieve. Strong free badges are reserved for tools whose access was actually checked.
79 tools found
It creates realistic custom voices from a few minutes of recorded speech and lets you control emotion and language.
It reads your text or PDFs aloud with natural voices and lets you download them as audio files.
It’s a free open-source lab that lets you make brand‑new music samples from text or existing audio.
It turns your music collection into a searchable sample library by splitting songs into stems and tagging them automatically.
It composes original songs in hundreds of styles and lets you edit them, even training the AI with your own influences.
It lets you instantly create custom background music tracks for your videos by selecting a mood and style.
Gemini 3.1 Flash TTS generates expressive multi-speaker speech with natural-language control over tone, pacing, and emotion.
Noiz Agent creates expressive cloned voices, multilingual dubbing, and long-form narration for podcasts, audiobooks, and video.
Convert text prompts into high-fidelity, royalty-free musical compositions for instant commercial use.
It helps you transform any sound in real time with Neutone's new Morpho model.
A high-fidelity neural voice synthesis engine offering sub-10-second cloning and a library of 300+ emotionally expressive multilingual voices.
Instantly transform written scripts into high-fidelity, natural-sounding audio in over 40 languages for professional media distribution.
A conversational production environment that transforms text-based creative direction into studio-grade musical stems and compositions.
An open-source generative audio model designed for high-fidelity musical composition and rapid soundscape arrangement.
Transform dense academic papers into engaging, conversational audio discussions using advanced generative synthesis.
Convert your static documents and articles into a dynamic, two-person conversational podcast for eyes-free learning.
It helps you make realistic and varied dialogs with a powerful text-to-speech model, accessible directly online.
It helps you get live transcripts with ultra-low latency of 150 ms, support for over 90 languages, word-for-word timestamps, and segments ready for u...