Mellotron
Mellotron: a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data
About
In our recent paper we propose Mellotron: a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data.
Open Source Health
- Stars
- 870
- Forks
- 187
- License
- BSD-3-Clause
- Last commit
- 3 years ago
Related Categories
Vendor
NVIDIA
Publisher of DALI (NVIDIA Data Loading Library) and Tensorrt
Quick Links
Open Source
More by NVIDIA
Related Products
say.js
TTS (Text To Speech) Module for Node.js
Top category match
Voicebox
The open-source AI voice studio. Clone, dictate, create.
Top category match
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
Top category match
OuteTTS
Interface for OuteTTS models.
Top category match
Qwen3-TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Top category match