Curator
Scalable data pre processing and curation toolkit for LLMs
About
NeMo Curator helps ML engineers and data teams build repeatable, GPU-accelerated pipelines that load, filter, deduplicate, and transform large text, image, video, and audio datasets for AI training. Run the same pipeline on a laptop or across a multi-node Ray cluster.
Open Source Health
- Stars
- 1,764
- Forks
- 323
- License
- Apache-2.0
- Last commit
- 29 days ago
Related Categories
Vendor
NVIDIA-NeMo
Open-source projects on GitHub: Speech, Guardrails and Switchyard
Quick Links
Open Source
More by NVIDIA-NeMo
Related Products
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
More from this vendor
Switchyard
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
More from this vendor
DataDesigner
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
More from this vendor
session.js
Session.js - Get user session information
Shared categories
Matrix
Matrix (Multi-Agent daTa geneRation Infra and eXperimentation framework) is a versatile engine for multi-agent conversational data generation.
Shared categories