Marketplace

Open Source Marketplace

Compare Open Source products across curated subcategories and trusted providers

Products
12,881
Subcategories
19

Selected subcategory

Voice, Speech & Audio AI

Compare Open Source products across curated subcategories and trusted providers

Explore Open Source with structured category paths, practical filters, and independent product data

Showing 1–12 of 30

Vosk Api

by Alpha Cephei

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

TranscriptionVoice, Speech & Audio AIOpen Source

Speech

by NVIDIA-NeMo

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

TranscriptionVoice, Speech & Audio AIOpen Source

OuteTTS

by edwko

Interface for OuteTTS models.

AI Voice SynthesisVoice, Speech & Audio AIOpen Source

Qwen3-TTS

by Qwenlm

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

AI Voice SynthesisVoice, Speech & Audio AIOpen Source

VoxCPM

by OpenBMB

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

AI Voice SynthesisVoice, Speech & Audio AIOpen Source

MOSS-Transcribe-Diarize

by OpenMOSS (SII)

A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness

TranscriptionVoice, Speech & Audio AIOpen Source

Sherpa Ncnn

by k2-fsa

Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

TranscriptionVoice, Speech & Audio AIOpen Source

MT3

by Magenta

MT3: Multi-Task Multitrack Music Transcription

TranscriptionVoice, Speech & Audio AIOpen Source

FireRedASR

by FireRedTeam

Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.

TranscriptionVoice, Speech & Audio AIOpen Source

Tacotron

by Kyubyong Park

A TensorFlow Implementation of Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model

AI Voice SynthesisVoice, Speech & Audio AIOpen Source

PyKaldi

by pykaldi

A Python wrapper for Kaldi

TranscriptionVoice, Speech & Audio AIOpen Source

FireRedASR2S

by FireRedTeam

A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID…

TranscriptionVoice, Speech & Audio AIOpen Source

Frequently asked questions