dots.tts.edit

A unified model for precise text and acoustic editing of speech.

dots.tts.edit follows natural, localized edit instructions while preserving the speaker and the acoustic context outside the edited regions. The same model supports text, emotion, prosody, pause, and compositional speech editing. Built on dots.tts.

Overview

dots.tts.edit system architecture

Evaluation

Task and language distribution in doteBench
Comparison of speech editing models
Loading audio examples…