Flowtron
Flowtron is an auto-regressive flow-based generative network for text to speech synthesis with control over speech variation and style transfer
About
In our recent paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer. Flowtron borrows insights from Autoregressive Flows and revamps Tacotron in order to provide high-quality and expressive mel-spectrogram synthesis. Flowtron is optimized by maximizing the likelihood of the training data, which makes training simple and stable. Flowtron learns an invertible mapping of data to a latent space that can be manipulated to control many aspects of speech synthesis (pitch, tone, speech rate, cadence, accent).
Open Source Health
- Stars
- 894
- Forks
- 175
- License
- Apache-2.0
- Last commit
- 3 years ago
Related Categories
Vendor
NVIDIA
Publisher of DALI (NVIDIA Data Loading Library) and Tensorrt
Quick Links
Open Source
More by NVIDIA
Related Products
say.js
TTS (Text To Speech) Module for Node.js
Top category match
Voicebox
The open-source AI voice studio. Clone, dictate, create.
Top category match
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
Top category match
OuteTTS
Interface for OuteTTS models.
Top category match
Qwen3-TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Top category match