这是一个从零开始构建的强化学习人类反馈(RLHF)学习代码库,实现了 PPO、GRPO、GSPO 以及相关的策略优化算法,并提供了清晰、可复现的训练流程。由于文档是由latex文件转译过来,如果md文件渲染异常,请用VScode的md插件打开
☆89Dec 19, 2025Updated 9 months ago
Alternatives and similar repositories for RLHF_learn
Users that are interested in RLHF_learn are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.☆171Mar 28, 2026Updated 6 months ago
- mini project for nanorllm☆63Mar 31, 2026Updated 6 months ago
- 多Agent金融研究报告自动生成系统 | Python/Java/Go三语言实现 | Pipeline+Fan-out并行架构 | 面试全套材料☆27Apr 6, 2026Updated 5 months ago
- code☆14Dec 9, 2024Updated last year
- 黑猫投诉平台,舆论监控系统☆20Apr 25, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。☆5,848Sep 15, 2026Updated 2 weeks ago
- ☆16Feb 23, 2025Updated last year
- Style Guide is a voice enabled AI assistant that lets you talk to your wardrobe. When you upload your photos, it analyzes them and identi…☆16Jul 20, 2024Updated 2 years ago
- ☆12Aug 10, 2023Updated 3 years ago
- [USENIX Security 2025] Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models☆19Jun 21, 2025Updated last year
- [ EMNLP 2025 Main ] Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs☆18Nov 7, 2025Updated 10 months ago
- Auto206Agent is an intelligent agent designed for research group scenarios, capable of supporting group document management as well as se…☆17May 29, 2026Updated 4 months ago
- Open Source Road Datasets☆19Aug 30, 2024Updated 2 years ago
- 🎯 Build a winning recommendation system with this effective generative framework, advancing to the finals of the 2025 Tencent Advertisin…☆27Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data (ACM MM '25)☆15Sep 12, 2025Updated last year
- ☆14Dec 9, 2024Updated last year
- ☆20Nov 23, 2025Updated 10 months ago
- ☆102May 4, 2026Updated 4 months ago
- Learn Microservices with Spring Boot (2nd edition) - Chapter 6☆23Feb 13, 2026Updated 7 months ago
- Edu-RAG:面向 K12 / 学科教育场景的轻量化 RAG 智能助教开源框架,依托教材、教辅知识库实现精准答疑、习题解析与个性化导学,有效抑制大模型知识幻觉,开箱即用快速落地 AI 教育应用☆53Jun 15, 2026Updated 3 months ago
- LW-CTrans☆18Jun 4, 2024Updated 2 years ago
- A project implementing various agentic RL based on the Slime post-training framework☆539Apr 11, 2026Updated 5 months ago
- 腾讯广告算法大赛2026-学术赛道-0.83220代码☆29May 24, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- AI驱动的竞品分析Agent协作系统 - 字节AI全栈挑战赛2026☆40May 16, 2026Updated 4 months ago
- ☆24Nov 4, 2025Updated 10 months ago
- LangGraph agent template with MCP.☆29Apr 8, 2025Updated last year
- Official code for paper "SPA-RL: Reinforcing LLM Agent via Stepwise Progress Attribution"☆93Sep 13, 2025Updated last year
- [AAAI 2026] Official implementation of "Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems".☆20Mar 23, 2026Updated 6 months ago
- 华南理工大学软件学院历年考试资料☆16Dec 6, 2021Updated 4 years ago
- ☆17Sep 3, 2025Updated last year
- Sharing a suggested AI learning roadmap, for reference only. Additions and suggestions are welcome.☆15Nov 8, 2025Updated 10 months ago
- Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.☆23Aug 28, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- AIR-Embodied: An Efficient Active 3DGS-based Interaction and Reconstruction Framework with Embodied Large Language Model☆22Apr 18, 2025Updated last year
- ☆398Aug 12, 2025Updated last year
- ☆54Oct 27, 2025Updated 11 months ago
- 手撕transformer并完成一个简单的机器翻译。☆23Feb 19, 2025Updated last year
- Official implementation for NeurIPS 2025 paper "SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tok…☆22Nov 7, 2025Updated 10 months ago
- 如何在github上传本地项目代码(新手使用)☆23Sep 30, 2019Updated 7 years ago
- Official Pytorch Implementation for "TextToucher: Fine-Grained Text-to-Touch Generation" (AAAI 2025)☆19Jan 28, 2026Updated 8 months ago