An open-source latent diffusion model that synthesizes complete songs with vocals and instrumentation in under ten seconds.
Search by what you want to achieve. Strong free badges are reserved for tools whose access was actually checked.
75 tools found
An open-source latent diffusion model that synthesizes complete songs with vocals and instrumentation in under ten seconds.
A high-precision foley synthesis engine that automatically synchronizes environmental and action-based audio with generative video sequences.
It helps you make personalized songs from your own lyrics or prompts (audio or text).
It helps you a free, open-source model capable of generating quality music in record time.
Zonos AI delivers a high-performance open-source alternative to proprietary speech APIs, enabling instantaneous voice cloning and expressive audio ge...
It helps you transcribe your audio files into text with remarkable accuracy thanks to this open source speech recognition model.
It helps you transform your texts and audio files into never-before-heard sounds with Fugatto from Nvidia.
Perform and evolve music in real-time using generative AI to bridge genres and manipulate sound textures through a responsive performance interface.
A visual-to-audio synthesis engine that transforms sketches and text into high-fidelity sound effects with granular acoustic control.
It helps you kokoro-82M TTS is an ultra-fast, open-source text-to-speech model that transforms your texts into natural speech (82M parameters only).
It helps you make dubbing, narration and dialogue with natural, expressive rendering.
It helps you easily transcribe medical dictations and doctor-patient conversations with Google's MedASR.
Alibaba's high-fidelity neural speech engine for natural, dialect-aware English and Chinese vocal synthesis via API.
Perform ultra-fast, local text-to-speech generation directly in the browser without server calls or GPU requirements.
Automate the creation of frame-accurate cinematic soundtracks and sound effects tailored to your visual timeline.
It helps you remove unwanted background noise from your audio files with Voice Isolator from ElevenLabs.
It helps you automatically transcribe and translate audio documents in over 50 languages with Whisper Large V3 Turbo.
An intelligent oratorical coach that provides real-time feedback and data-driven analytics to refine public speaking performance.