LLM OrchestrationOpen Source AI & Machine Learning

MOSS-TTSD

A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning

Open Source

About

MOSS-TTSD is the long-form dialogue specialist within our open-source MOSS‑TTS Family. While foundational models typically prioritize high-fidelity single-speaker synthesis, MOSS-TTSD is architected to bridge the gap between isolated audio samples and cohesive, continuous human interaction. The model represents a paradigm shift from "text-to-speech" to "script-to-conversation." By prioritizing the flow and emotional nuances of multi-party engagement, MOSS-TTSD transforms static dialogue scripts into dynamic, expressive oral performances. It is designed to serve as a robust backbone for creators and developers who require a seamless transition between distinct speaker personas without sacrificing narrative continuity. Whether it is capturing the spontaneous energy of a live talk show or the structured complexity of a multilingual drama, MOSS-TTSD provides the stability and expressive depth necessary for professional-grade, long-form content creation in an open-source framework.

Open Source Health

Not enough history
Stars
1,395
Forks
137
License
Apache-2.0
Last commit
1 months ago
Python

Related Categories