这是一个从零开始构建的强化学习人类反馈(RLHF)学习代码库,实现了 PPO、GRPO、GSPO 以及相关的策略优化算法,并提供了清晰、可复现的训练流程。由于文档是由latex文件转译过来,如果md文件渲染异常,请用VScode的md插件打开
☆89Dec 19, 2025Updated 8 months ago
Alternatives and similar repositories for RLHF_learn
Users that are interested in RLHF_learn are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is the official code for OThink-R1 project.☆21Jun 19, 2025Updated last year
- A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.☆170Mar 28, 2026Updated 5 months ago
- generate B-scan for training neural network☆11Dec 8, 2022Updated 3 years ago
- mini project for nanorllm☆63Mar 31, 2026Updated 5 months ago
- 多Agent金融研究报告自动生成系统 | Python/Java/Go三语言实现 | Pipeline+Fan-out并行架构 | 面试全套材料☆23Apr 6, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- code☆14Dec 9, 2024Updated last year
- 本项目是一个基于LangChain构建的多Agent系统,结合Streamlit实现的Web界面,能够根据用户输入进行网络搜索并提供旅游相关的聊天服务。此外,该系统还具备基于本地知识库的推销功能,为用户提供个性化的旅游产品推荐。☆16Apr 20, 2025Updated last year
- Code for "Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models", ICLR 2024 Oral.☆21Feb 4, 2026Updated 7 months ago
- collection with description of super-resolution related papers, repositories, datasets, loss functions and etc.☆11Dec 12, 2023Updated 2 years ago
- 个人整理的AI PM资料库,沉 淀多年AI产品管理经验与方法论。包含精选工具、实用模板和踩坑总结,涵盖需求分析、模型评估到落地交付全流程。附有个人实践中的案例复盘和心得思考。持续更新在AI产品一线的实战干货,欢迎交流讨论,共同成长。☆16Jul 23, 2025Updated last year
- MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。☆5,799Updated this week
- ☆16Feb 23, 2025Updated last year
- ☆12Sep 30, 2024Updated last year
- Style Guide is a voice enabled AI assistant that lets you talk to your wardrobe. When you upload your photos, it analyzes them and identi…☆16Jul 20, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official implementation of "PersonaBooth: Personalized Text-to-Motion Generation (CVPR 2025)"☆37Sep 27, 2025Updated 11 months ago
- Auto206Agent is an intelligent agent designed for research group scenarios, capable of supporting group document management as well as se…☆18May 29, 2026Updated 3 months ago
- Open Source Road Datasets☆19Aug 30, 2024Updated 2 years ago
- ☆23Mar 20, 2024Updated 2 years ago
- 🎯 Build a winning recommendation system with this effective generative framework, advancing to the finals of the 2025 Tencent Advertisin…☆27Updated this week
- ☆14Dec 9, 2024Updated last year
- ☆99May 4, 2026Updated 4 months ago
- 中文工业技术文档多模态推理问答数据集☆16Sep 23, 2025Updated 11 months ago
- A project implementing various agentic RL based on the Slime post-training framework☆533Apr 11, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 腾讯广告算法大赛2026-学术赛道-0.83220代码☆28May 24, 2026Updated 3 months ago
- AI驱动的竞品分析Agent协作系统 - 字节AI全栈挑战赛2026☆37May 16, 2026Updated 3 months ago
- [ACL 2024 Findings] The code for Beyond Single-Event Extraction: Towards Efficient Document-Level Multi-Event Argument Extraction☆23Nov 4, 2024Updated last year
- ☆35Aug 19, 2026Updated 3 weeks ago
- Official repo for SAO-Instruct: Free-form Audio Editing using Natural Language Instructions presented at NeurIPS 2025☆19Oct 28, 2025Updated 10 months ago
- ☆24Nov 4, 2025Updated 10 months ago
- My implementation in TianChi CCKS 2025 pdf QA multimodal competition☆19Aug 27, 2025Updated last year
- This repository contains the official implementation of the paper "LandSegmenter: Towards a Flexible Foundation Model for Land Use and La…☆45Jul 13, 2026Updated 2 months ago
- [AAAI 2026] Official implementation of "Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems".☆19Mar 23, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR'26] "Nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" by Peihao Wang*, Ruisi Cai*, Zhen Wang, Hongyuan…☆36Mar 10, 2026Updated 6 months ago
- 华南理工大学软件学院历年考试资料☆16Dec 6, 2021Updated 4 years ago
- ☆17Sep 3, 2025Updated last year
- Frequency Shortcuts in Neural Networks☆21Nov 1, 2024Updated last year
- This is the source code for our accepted paper in the journal RA_L with IROS2021.☆22Jul 8, 2021Updated 5 years ago
- Diffusion Model-Augmented Behavioral Cloning☆22Oct 8, 2024Updated last year
- AIR-Embodied: An Efficient Active 3DGS-based Interaction and Reconstruction Framework with Embodied Large Language Model☆22Apr 18, 2025Updated last year