Marketplace
Open Source Marketplace
Compare Open Source products across curated subcategories and trusted providers
- Products
- 12,881
- Subcategories
- 19
Selected subcategory
Voice, Speech & Audio AI
Compare Open Source products across curated subcategories and trusted providers
Explore Open Source with structured category paths, practical filters, and independent product data
Showing 1–12 of 30
by Alpha Cephei
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
by NVIDIA-NeMo
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
by Qwenlm
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
by OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
by OpenMOSS (SII)
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
by k2-fsa
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.
by Magenta
MT3: Multi-Task Multitrack Music Transcription
by FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
by Kyubyong Park
A TensorFlow Implementation of Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model
by FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID…