AuK
AuK: An Open-Source Foundational Model for Speech Generation and Editing
About
AuK is a foundational model that focuses on various aspects of speech technology, including speaker extraction, speech enhancement, and text-to-speech capabilities. Built using Python and PyTorch, it supports deep learning applications in speech editing and generation. This model is suitable for developers and researchers interested in advancing speech-related technologies.
Open Source Health
- Stars
- 708
- Forks
- 48
- License
- Not stated
- Last commit
- 27 days ago
Related Categories
Vendor
Tencent-Hunyuan
Open-source projects on GitHub: HunyuanVideo, HunyuanVideo-1.5 and HunyuanOCR
Quick Links
Open Source
More by Tencent-Hunyuan
Related Products
Rclip
Semantic photo search for the command line
Top category match
ChineseNER
A neural network model for Chinese named entity recognition
Top category match
Head Pose Estimation
Realtime human head pose estimation with ONNXRuntime and OpenCV.
Top category match
Video Subtitle Extractor
视频硬字幕提取,生成srt文件。无需申请第三方API,本地实现文本识别。基于深度学习的视频字幕提取框架,包含字幕区域检测、字幕内容提取。A GUI tool for extracting hard-coded subtitle (hardsub) from videos and generating srt files.
Top category match
U 2 Net
The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection."
Top category match