It creates realistic custom voices from a few minutes of recorded speech and lets you control emotion and language.
Audio & Voice

It clones any voice from a three-second audio sample and can speak new text in that voice with emotion.
Best for: voice content creators, accessibility developers and linguists
It creates realistic custom voices from a few minutes of recorded speech and lets you control emotion and language.
It turns your text into high-quality speech and lets you customise voice profiles, speed and pitch.
A high-performance open-source text-to-speech model providing granular control over vocal characteristics and custom cloning.