Mattral / Improving-LLM-Models-with-RLHF-PPO-DPOView on GitHub
A modular, production-grade framework for Reinforcement Learning from Human Feedback (RLHF) with Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO).
23Jun 6, 2026Updated last month

Alternatives and similar repositories for Improving-LLM-Models-with-RLHF-PPO-DPO

Users that are interested in Improving-LLM-Models-with-RLHF-PPO-DPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?