Tacotron
A TensorFlow Implementation of Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model
About
LJ Speech Dataset is recently widely used as a benchmark dataset in the TTS task because it is publicly available. It has 24 hours of reasonable quality samples. Nick's audiobooks are additionally used to see if the model can learn even with less data, variable speech samples. They are 18 hours long. The World English Bible is a public domain update of the American Standard Version of 1901 into modern English. Its original audios are freely available here. Kyubyong split each chapter by verse manually and aligned the segmented audio clips to the text. They are 72 hours in total. You can download them at Kaggle Datasets.
Open Source Health
- Stars
- 1,832
- Forks
- 427
- License
- Apache-2.0
- Last commit
- 5 years ago
Related Categories
Vendor
Kyubyong Park
Lives in Seoul, Korea. Studied Linguistics at SNU and Univ. of Hawaii
Quick Links
Open Source
More by Kyubyong Park
Related Products
say.js
TTS (Text To Speech) Module for Node.js
Top category match
Voicebox
The open-source AI voice studio. Clone, dictate, create.
Top category match
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
Top category match
OuteTTS
Interface for OuteTTS models.
Top category match
Qwen3-TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Top category match