Train transformer language models with reinforcement learning.
β19,443Oct 3, 2026Updated this week
Alternatives and similar repositories for trl
Users that are interested in trl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,749Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Frameworkβ23,738Updated this week
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asyβ¦β10,066Sep 17, 2026Updated 2 weeks ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)β4,755Jan 8, 2024Updated 2 years ago
- A modular RL library to fine-tune language models to human preferencesβ2,396Mar 1, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fast and memory-efficient exact attentionβ25,068Updated this week
- Fully open reproduction of DeepSeek-R1β26,478Updated this week
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)β75,291Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β43,172Updated this week
- Robust recipes to align language models with human and AI preferencesβ5,686Sep 23, 2026Updated last week
- A high-throughput and memory-efficient inference and serving engine for LLMsβ93,111Updated this week
- Ongoing research training transformer models at scaleβ18,061Updated this week
- A framework for few-shot evaluation of language models.β14,125Sep 14, 2026Updated 2 weeks ago
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,555May 1, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Large Language Model Text Generation Inferenceβ10,884Mar 21, 2026Updated 6 months ago
- Reference implementation for DPO (Direct Preference Optimization)β2,911Aug 11, 2024Updated 2 years ago
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,899Updated this week
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,234Jul 17, 2024Updated 2 years ago
- [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.β25,057Aug 12, 2024Updated 2 years ago
- Example models using DeepSpeedβ6,849Updated this week
- QLoRA: Efficient Finetuning of Quantized LLMsβ11,030Jun 10, 2024Updated 2 years ago
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VLβ¦β15,776Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.β36,748Updated this week
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,226Sep 21, 2026Updated last week
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β166,917Updated this week
- Accessible large language models via k-bit quantization for PyTorch.β8,510Sep 7, 2026Updated 3 weeks ago
- Inference code for Llama modelsβ59,619Jan 26, 2025Updated last year
- A curated list of reinforcement learning with human feedback resources (continually updated)β4,431May 20, 2026Updated 4 months ago
- Instruct-tune LLaMA on consumer hardwareβ18,900Jul 29, 2024Updated 2 years ago
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.β77,167Updated this week
- Democratizing Reinforcement Learning for LLMsβ5,856Sep 12, 2026Updated 3 weeks ago
- slime is an LLM post-training framework for RL Scaling.β8,589Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Aligning pretrained language models with instruction data generated by themselves.β4,612Mar 27, 2023Updated 3 years ago
- Tools for merging pretrained large language models.β7,386Sep 12, 2026Updated 3 weeks ago
- Simple RL training for reasoningβ3,875Dec 23, 2025Updated 9 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRLβ5,182Sep 19, 2026Updated 2 weeks ago
- Code for the paper Fine-Tuning Language Models from Human Preferencesβ1,388Jul 25, 2023Updated 3 years ago
- Making large AI models cheaper, faster and more accessibleβ41,441Updated this week
- Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We alsβ¦β18,557May 19, 2026Updated 4 months ago