Notes and commented code for RLHF (PPO)
☆136Feb 27, 2024Updated 2 years ago
Alternatives and similar repositories for rlhf-ppo
Users that are interested in rlhf-ppo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Notes on Direct Preference Optimization☆28Apr 14, 2024Updated 2 years ago
- Reference implementation of Mistral AI 7B v0.1 model.☆28Dec 25, 2023Updated 2 years ago
- LLaMA 2 implemented from scratch in PyTorch☆375Sep 25, 2023Updated 2 years ago
- Distributed training (multi-node) of a Transformer model☆98Apr 10, 2024Updated 2 years ago
- Because it's there.☆16Sep 22, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official repository for our paper, Transformers Learn Higher-Order Optimization Methods for In-Context Learning: A Study with Linear Mode…☆20Nov 19, 2024Updated last year
- An unnecessarily tiny and minimal implementation of GPT-2 in NumPy.☆11Feb 12, 2023Updated 3 years ago
- Training codebase for K2-V2☆22Dec 17, 2025Updated 7 months ago
- Notes on the Mistral AI model