Tacotron

A TensorFlow Implementation of Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model

AI Voice SynthesisVoice, Speech & Audio AIOpen Source

About

LJ Speech Dataset is recently widely used as a benchmark dataset in the TTS task because it is publicly available. It has 24 hours of reasonable quality samples. Nick's audiobooks are additionally used to see if the model can learn even with less data, variable speech samples. They are 18 hours long. The World English Bible is a public domain update of the American Standard Version of 1901 into modern English. Its original audios are freely available here. Kyubyong split each chapter by verse manually and aligned the segmented audio clips to the text. They are 72 hours in total. You can download them at Kaggle Datasets.

Open Source Health

Not enough history
Stars
1,832
Forks
427
License
Apache-2.0
Last commit
5 years ago
Python

Resources & Links

Related Categories

Using Tacotron?

Track its cost next to the rest of your stack and get a reminder before it renews.

Add to my stack

Vendor

Kyubyong Park

Lives in Seoul, Korea. Studied Linguistics at SNU and Univ. of Hawaii

View vendor profile

More by Kyubyong Park