It turns your text into high-quality speech and lets you customise voice profiles, speed and pitch.
Audio & Voice

Alibaba's high-fidelity neural speech engine for natural, dialect-aware English and Chinese vocal synthesis via API.
Best for: Large-scale enterprise localization and high-fidelity narration for Mandarin-English bimodal applications.
It turns your text into high-quality speech and lets you customise voice profiles, speed and pitch.
It helps you kokoro-82M TTS is an ultra-fast, open-source text-to-speech model that transforms your texts into natural speech (82M parameters only).
Perform ultra-fast, local text-to-speech generation directly in the browser without server calls or GPU requirements.