Slime
slime is an LLM post-training framework for RL Scaling.
About
slime's design goal is to make these two capabilities reinforce each other without turning the system into a heavy stack of disconnected trainers, rollout services, and agent frameworks. Megatron training, SGLang rollout, custom data generation, reward computation, verifier feedback, and environment interaction all flow through the same training / rollout / Data Buffer path.
Open Source Health
- Stars
- 8,457
- Forks
- 1,249
- License
- Apache-2.0
- Last commit
- 1 months ago
Resources & Links
Related Categories
Vendor
THUKEG
Open-source projects on GitHub: Slime, AgentBench and P-tuning-v2
Quick Links
Open Source
More by THUKEG
Related Products
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories
Textrank
TextRank implementation for Python 3.
Shared categories