SpeechT5
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
About
Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-supervised speech/text representation learning. The SpeechT5 framework consists of a shared encoder-decoder network and six modal-specific (speech/text) pre/post-nets. After preprocessing the input speech/text through the pre-nets, the shared encoder-decoder network models the sequence-to-sequence transformation, and then the post-nets generate the output in the speech/text modality based on the output of the decoder.
Open Source Health
- Stars
- 1,450
- Forks
- 135
- License
- MIT
- Last commit
- 2 years ago
Related Categories
Vendor
Microsoft
Empower every person and organization to achieve more.
Quick Links
Open Source
More by Microsoft
Related Products
Whisper Ctranslate2
Whisper command line client compatible with original OpenAI client based on CTranslate2.
Top category match
Voxtype
Voice-to-text with push-to-talk for Wayland compositors
Top category match
pyVideoTrans
Translate the video from one language to another and embed dubbing & subtitles.
Top category match
Kaldi Gstreamer Server
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
Top category match
Vosk Api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Top category match
