Tacotron
A TensorFlow Implementation of Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model
About
LJ Speech Dataset is recently widely used as a benchmark dataset in the TTS task because it is publicly available. It has 24 hours of reasonable quality samples. Nick's audiobooks are additionally used to see if the model can learn even with less data, variable speech samples. They are 18 hours long. The World English Bible is a public domain update of the American Standard Version of 1901 into modern English. Its original audios are freely available here. Kyubyong split each chapter by verse manually and aligned the segmented audio clips to the text. They are 72 hours in total. You can download them at Kaggle Datasets.
Open Source Health
- Stars
- 1,832
- Forks
- 427
- License
- Apache-2.0
- Last commit
- 5 years ago
Alternatives to Tacotron
say.jsby MarakTTS (Text To Speech) Module for Node.js
Voiceboxby Jamie PineThe open-source AI voice studio. Clone, dictate, create.
OpenVoiceby MyShellInstant voice cloning by MIT and MyShell. Audio foundation model.
OuteTTSby edwkoInterface for OuteTTS models.
Qwen3-TTSby QwenlmQwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Resources & Links
Related Categories
Using Tacotron?
Track its cost next to the rest of your stack and get a reminder before it renews.
Add to my stackVendor
Kyubyong Park
Lives in Seoul, Korea. Studied Linguistics at SNU and Univ. of Hawaii