Fun-Audio-Chat
Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions.
About
Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions. It introduces Dual-Resolution Speech Representations (an efficient 5Hz shared backbone + a 25Hz refined head) to cut compute while keeping high speech quality, and Core-Cocktail training to preserve strong text LLM capabilities. It delivers top-tier results on spoken QA, audio understanding, speech function calling, and speech instruction-following and voice empathy benchmarks.
Open Source Health
- Stars
- 1,004
- Forks
- 105
- License
- Apache-2.0
- Last commit
- 8 months ago
Resources & Links
Related Categories
Vendor
QwenAudio
Open-source speech and audio language models from the QwenAudio Team
Quick Links
Open Source
More by QwenAudio
Related Products
Dogehouse
Taking voice conversations to the moon 🚀
Top category match
Figaro
Real-time voice-changer for voice-chat, etc. Will support many different voice-filters and features in the future. 🎵
Top category match
CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
More from this vendor
Qwen Audio Agent
A realtime voice frontend for mainstream AI agents.
More from this vendor
SenseVoice
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
More from this vendor