LongAlign
[EMNLP 2024] LongAlign: A Recipe for Long Context Alignment of LLMs
About
LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on training strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluates the instruction-following capability on queries of 10k-100k length.
Open Source Health
- Stars
- 262
- Forks
- 21
- License
- Apache-2.0
- Last commit
- 2 years ago
Related Categories
Vendor
THUKEG
Open-source projects on GitHub: Slime, AgentBench and P-tuning-v2
Quick Links
Open Source
More by THUKEG
Related Products
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories
Textrank
TextRank implementation for Python 3.
Shared categories