Train transformer language models with reinforcement learning.
β19,002Aug 4, 2026Updated this week
Alternatives and similar repositories for trl
Users that are interested in trl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,501Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Frameworkβ22,797Updated this week
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asyβ¦β9,882Jul 14, 2026Updated 3 weeks ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)β4,752Jan 8, 2024Updated 2 years ago
- A modular RL library to fine-tune language models to human preferencesβ2,394Mar 1, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Fast and memory-efficient exact attentionβ24,620Updated this week
- Fully open reproduction of DeepSeek-R1β26,422Apr 2, 2026Updated 4 months ago
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)β73,749Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β42,858Updated this week
- Robust recipes to align language models with human and AI preferencesβ5,653May 26, 2026Updated 2 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMsβ88,177Updated this week
- Ongoing research training transformer models at scaleβ17,320Updated this week
- A framework for few-shot evaluation of language models.β13,528Jul 13, 2026Updated 3 weeks ago
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,514May 1, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Large Language Model Text Generation Inferenceβ10,887Mar 21, 2026Updated 4 months ago
- Reference implementation for DPO (Direct Preference Optimization)β2,900Aug 11, 2024Updated last year
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,807Updated this week
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,245Jul 17, 2024Updated 2 years ago
- [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.β24,965Aug 12, 2024Updated last year
- Example models using DeepSpeedβ6,833Updated this week
- QLoRA: Efficient Finetuning of Quantized LLMsβ10,982Jun 10, 2024Updated 2 years ago
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VLβ¦β15,045Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.β31,270Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,178Jan 23, 2026Updated 6 months ago
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β163,335Updated this week
- Accessible large language models via k-bit quantization for PyTorch.β8,387Jul 29, 2026Updated last week
- Inference code for Llama modelsβ59,538Jan 26, 2025Updated last year
- A curated list of reinforcement learning with human feedback resources (continually updated)β4,421May 20, 2026Updated 2 months ago
- Instruct-tune LLaMA on consumer hardwareβ18,912Jul 29, 2024Updated 2 years ago
- Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.β69,555Updated this week
- Democratizing Reinforcement Learning for LLMsβ5,759Updated this week
- slime is an LLM post-training framework for RL Scaling.β7,758Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tools for merging pretrained large language models.β7,281Jun 17, 2026Updated last month
- Aligning pretrained language models with instruction data generated by themselves.β4,608Mar 27, 2023Updated 3 years ago
- Simple RL training for reasoningβ3,873Dec 23, 2025Updated 7 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRLβ5,099Updated this week
- Code for the paper Fine-Tuning Language Models from Human Preferencesβ1,393Jul 25, 2023Updated 3 years ago
- Making large AI models cheaper, faster and more accessibleβ41,432Jul 13, 2026Updated 3 weeks ago
- Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We alsβ¦β18,551May 19, 2026Updated 2 months ago