dots.tts.edit

A unified model for precise text and acoustic editing of speech.

Interactive Playground Paper

dots.tts.edit follows natural, localized edit instructions while preserving the speaker and the acoustic context outside the edited regions. The same model supports text, emotion, prosody, pause, and compositional speech editing.

Contents

  • Overview
  • Evaluation
  • Text Editing
  • Emotion Editing
  • Prosody Editing
  • Pause Editing
  • Compositional Editing

Overview

dots.tts.edit system architecture

Evaluation

Task and language distribution in doteBench
Comparison of speech editing models
Loading audio examples…

dots.tts.edit · Research demo

Audio samples are provided for research demonstration only. Clearly identify AI-generated audio.