SenseVoice
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
About
SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).
Open Source Health
- Stars
- 9,297
- Forks
- 823
- License
- MIT
- Last commit
- 29 days ago
Related Categories
Vendor
QwenAudio
Open-source speech and audio language models from the QwenAudio Team
Quick Links
Open Source
More by QwenAudio
Related Products
Whisper Ctranslate2
Whisper command line client compatible with original OpenAI client based on CTranslate2.
Top category match
Voxtype
Voice-to-text with push-to-talk for Wayland compositors
Top category match
pyVideoTrans
Translate the video from one language to another and embed dubbing & subtitles.
Top category match
Kaldi Gstreamer Server
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
Top category match
Vosk Api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Top category match