Turn English speech into punctuated text with word timestamps using an NVIDIA open ASR model.
Audio & Voice
Turn a video, social link, or audio file into a transcript, subtitles, and a summary.
Best for: Students, creators, podcasters, and teams turning recorded content into searchable text.
Turn English speech into punctuated text with word timestamps using an NVIDIA open ASR model.
A browser audio toolkit for text-to-speech, speech-to-text, vocal removal, voice enhancement, and quick editing.
Generate expressive speech and clone an authorized voice locally with an efficient open-source TTS model.