RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooks
☆253Jun 20, 2025Updated last year
Alternatives and similar repositories for RLHF_in_notebooks
Users that are interested in RLHF_in_notebooks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [experimental] multiplexed distributed tensor framework☆22Nov 17, 2025Updated 9 months ago
- Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.☆630Feb 24, 2025Updated last year
- A reimplementation of Stable Diffusion 3.5 in pure PyTorch☆707Jun 14, 2025Updated last year
- A theoretical and practical deep dive into Reinforcement Learning with Human Feedback and it’s applications in Large Language Models from…☆117Nov 7, 2025Updated 10 months ago
- Implementations of Papers that I read, you can read my breakdown in my blog☆91Oct 23, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆565Jul 1, 2025Updated last year
- Sample Random GitHub Repositories☆14Aug 31, 2026Updated last week
- Implement a reasoning LLM in PyTorch from scratch, step by step☆5,171Updated this week
- ☆13Aug 12, 2024Updated 2 years ago
- Static suckless single batch CUDA-only qwen3-0.6B mini inference engine☆559Sep 8, 2025Updated 11 months ago
- Thorn in a HaizeStack test for evaluating long-context adversarial robustness.☆26Aug 3, 2024Updated 2 years ago
- A JPEG Image Compression Service using Part Homomorphic Encryption.☆31Mar 7, 2025Updated last year
- ☆51Sep 28, 2025Updated 11 months ago
- A Python tool to parse PDF statements from Poste Italiane (Postepay, BancoPosta) and extract data as structured JSON.☆50Jul 25, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and full…☆641Mar 23, 2025Updated last year
- A plugin for LLM that enables proper plugin management when installed via uv tool☆35Jul 26, 2026Updated last month
- A powerful and user-friendly tool that generates detailed captions for your images☆21Nov 11, 2024Updated last year
- Temporal Graph Rewiring Method with Expander Graphs☆12Oct 18, 2024Updated last year
- Auto Thinking Mode switch for Qwen3 in Open webui☆71May 8, 2025Updated last year
- Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought☆20Feb 10, 2026Updated 6 months ago
- Cell type annotation with local Large Language Models (LLMs) - Ensuring privacy and speed with extensive customized reports☆153Oct 25, 2024Updated last year
- Probability and Statistics for Data Science: A self-contained introduction to probability and statistics for data science, including a fr…☆620Aug 27, 2026Updated last week
- A multi-agent system trained with GRPO for reliable long-horizon task planning and execution.☆61Feb 9, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 🧠 Web AI / LLM in browser / Whisper in browser / WebGPU inference Examples☆35Oct 1, 2025Updated 11 months ago
- Experimental implementation of DeepSeek v4 flaash in llama.cpp☆24Apr 30, 2026Updated 4 months ago
- A single-page webapp that decrypts text using only client-side JavaScript☆16Feb 9, 2023Updated 3 years ago
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 11 months ago
- 250+ Fine-tuning & RL Notebooks for text, vision, audio, embedding, TTS models.☆5,660Updated this week
- This is a framework that implements various parallel reasoning strategies from the literature☆275Dec 18, 2025Updated 8 months ago
- Build an email assistant with human-in-the-loop and memory☆2,181Aug 11, 2026Updated 3 weeks ago
- Inbound email processing server with LLM-powered automation☆36Aug 5, 2026Updated last month
- Open-source VMs-as-a-service☆799Aug 30, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Bring your dotfiles to remote machine via SSH☆172Feb 28, 2026Updated 6 months ago
- k for BareMetal☆13Dec 10, 2024Updated last year
- Ultra-lightweight AI Agent☆444Mar 5, 2026Updated 6 months ago
- Intercept LLM API traffic and visualize token usage in a real-time terminal dashboard. Track costs, debug prompts, and monitor cont…☆813Jun 21, 2026Updated 2 months ago
- ☆212Jan 5, 2026Updated 8 months ago
- A TV-Guide for hadisafa's YTCH.xyz via Hacker News (https://news.ycombinator.com/item?id=41247023)☆14Aug 26, 2024Updated 2 years ago
- A list of projects that enables you to use GPT-3.5 for free☆38May 8, 2024Updated 2 years ago