JustRL
[ICLR 2026 Blogpost Track Poster] JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
About
JustRL demonstrates that competitive reinforcement learning performance for small language models doesn't require complex multi-stage pipelines or dynamic schedules. Using a minimal recipe with single-stage training and fixed hyperparameters, we achieve state-of-the-art results on mathematical reasoning tasks. This repository contains a lightweight evaluation script to reproduce evaluation results for JustRL models on nine challenging math benchmarks.
Open Source Health
- Stars
- 305
- Forks
- 15
- License
- Not stated
- Last commit
- 3 months ago
Resources & Links
Related Categories
Vendor
THUNLP
Natural Language Processing Lab at Tsinghua University
Quick Links
Open Source
More by THUNLP
Related Products
TAADpapers
Must-read Papers on Textual Adversarial Attack and Defense
More from this vendor
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories