Train transformer language models with reinforcement learning.
β19,144Aug 24, 2026Updated this week
Alternatives and similar repositories for trl
Users that are interested in trl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,590Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Frameworkβ23,106Updated this week
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asyβ¦β9,950Aug 13, 2026Updated last week
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)β4,754Jan 8, 2024Updated 2 years ago
- A modular RL library to fine-tune language models to human preferencesβ2,394Mar 1, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Fast and memory-efficient exact attentionβ24,769Updated this week
- Fully open reproduction of DeepSeek-R1β26,443Apr 2, 2026Updated 4 months ago
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)β74,313Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β42,985Updated this week
- Robust recipes to align language models with human and AI preferencesβ5,665May 26, 2026Updated 2 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMsβ89,879Updated this week
- Ongoing research training transformer models at scaleβ17,563Updated this week
- A framework for few-shot evaluation of language models.β13,774Updated this week
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,524May 1, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Large Language Model Text Generation Inferenceβ10,887Mar 21, 2026Updated 5 months ago
- Reference implementation for DPO (Direct Preference Optimization)β2,906Aug 11, 2024Updated 2 years ago
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,831Updated this week
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,246Jul 17, 2024Updated 2 years ago
- [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.β25,002Aug 12, 2024Updated 2 years ago
- Example models using DeepSpeedβ6,841Updated this week
- QLoRA: Efficient Finetuning of Quantized LLMsβ10,997Jun 10, 2024Updated 2 years ago
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VLβ¦β15,343Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.β32,362Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,194Jan 23, 2026Updated 7 months ago
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β164,394Updated this week
- Accessible large language models via k-bit quantization for PyTorch.β8,436Updated this week
- Inference code for Llama modelsβ59,573Jan 26, 2025Updated last year
- A curated list of reinforcement learning with human feedback resources (continually updated)β4,423May 20, 2026Updated 3 months ago
- Instruct-tune LLaMA on consumer hardwareβ18,907Jul 29, 2024Updated 2 years ago
- Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.β74,594Updated this week
- Democratizing Reinforcement Learning for LLMsβ5,797Updated this week
- slime is an LLM post-training framework for RL Scaling.β8,231Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Tools for merging pretrained large language models.β7,305Jun 17, 2026Updated 2 months ago
- Aligning pretrained language models with instruction data generated by themselves.β4,610Mar 27, 2023Updated 3 years ago
- Simple RL training for reasoningβ3,874Dec 23, 2025Updated 8 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRLβ5,128Jul 30, 2026Updated 3 weeks ago
- Code for the paper Fine-Tuning Language Models from Human Preferencesβ1,391Jul 25, 2023Updated 3 years ago
- Making large AI models cheaper, faster and more accessibleβ41,438Aug 17, 2026Updated last week
- Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We alsβ¦β18,561May 19, 2026Updated 3 months ago