Simple-TTS: End-to-End Text-to-Speech Synthesis with Latent Diffusion
Justin Lovelace, sray@asapp.com, Kwangyoun Kim, Kilian Q Weinberger, Felix Wu
OpenReview ground truth
TL;DR — We introduce an end-to-end text-to-speech (TTS) latent diffusion model as a simpler alternative to more complicated pipelined approaches for TTS synthesis.
Abstract
We propose an end-to-end text-to-speech (TTS) latent diffusion model as a simpler alternative to more complicated pipelined approaches for TTS synthesis. In particular, we show that one can adapt a recently proposed text-to-image diffusion architecture, U-ViT, as an excellent backbone for audio generation. We identify and explain the changes required for this adaptation and demonstrate that latent diffusion is an effective approach for end-to-end speech synthesis, without the need for phonemizers, forced aligners, or complex multi-stage pipelines. Despite its simplicity, our proposed approach, Simple-TTS, outperforms more complex models that rely on explicit alignment components and significantly outperforms the best open-source multi-speaker TTS system. We will open-source Simple-TTS upon acceptance, making it the strongest system publicly available to the community. Due to its straight-forward design, we expect that Simple-TTS can easily be adapted to many diverse TTS settings --- opening the stage to repeat the success of Stable Diffusion in computer vision, in audio generation.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 59% of matchups.
- ▼ lost to Multimodal Molecular Pretraining via Modal… ×6
- ▼ lost to On Representation Complexity of Model-base… ×6
- ▼ lost to Bridging the gap between offline and onlin… ×6
- ▼ lost to In-Context Learning Learns Label Relations… ×6
- ▲ beat Improving Compositional Text-to-image Gene… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)