Jamify
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
About
JAM is a rectified flow-based model for lyrics-to-song generation that addresses the lack of fine-grained word-level controllability in existing lyrics-to-song models. Built on a compact 530M-parameter architecture with 16 LLaMA-style Transformer layers as the Diffusion Transformer (DiT) backbone, JAM enables precise vocal control that musicians desire in their workflows. Unlike previous models, JAM provides word and phoneme-level timing control, allowing musicians to specify the exact placement of each vocal sound for improved rhythmic flexibility and expressive timing.
Open Source Health
- Stars
- 169
- Forks
- 19
- License
- Not stated
- Last commit
- 1 years ago
Related Categories
Vendor
Deep Cognition and Language Research (DeCLaRe) Lab
Open-source projects on GitHub: Conv Emotion, Tango and MELD
Quick Links
Open Source
More by Deep Cognition and Language Research (DeCLaRe) Lab
Related Products
say.js
TTS (Text To Speech) Module for Node.js
Top category match
Voicebox
The open-source AI voice studio. Clone, dictate, create.
Top category match
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
Top category match
OuteTTS
Interface for OuteTTS models.
Top category match
Qwen3-TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Top category match