TranscriptionVoice, Speech & Audio AI

SpeechT5

Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

Open Source

About

Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-supervised speech/text representation learning. The SpeechT5 framework consists of a shared encoder-decoder network and six modal-specific (speech/text) pre/post-nets. After preprocessing the input speech/text through the pre-nets, the shared encoder-decoder network models the sequence-to-sequence transformation, and then the post-nets generate the output in the speech/text modality based on the output of the decoder.

Open Source Health

Not enough history
Stars
1,450
Forks
135
License
MIT
Last commit
2 years ago
Python

Related Categories