☆108Jul 24, 2025Updated last year
Alternatives and similar repositories for study_rlhf
Users that are interested in study_rlhf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆171Mar 18, 2026Updated 4 months ago
- MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。☆5,694Jun 3, 2026Updated 2 months ago
- Enhanced Search-R1 Implementation: Improved Compatibility and Modern Framework Integration☆31Dec 8, 2025Updated 8 months ago
- 复现大模型相关算法及一些学习记录☆3,492Jul 2, 2026Updated last month
- WebResearcher: An Iterative Deep-Research Agent,迭代式深度研究智能体☆50Feb 13, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 面向个人投资者金融投资研究与辅助决策agent系统☆29Jul 9, 2026Updated last month
- llm & rl☆293Oct 24, 2025Updated 9 months ago
- 🏥 从零基础到面试通关:20节课彻底搞懂MedicalGPT医疗大模型训练全流程 | PT/SFT/LoRA/RLHF/DPO/GRPO | 100+面试高频考点☆199Apr 1, 2026Updated 4 months ago
- ☆44Nov 22, 2025Updated 8 months ago
- Beyond Basic RAG, Empowering Real-Time Deep Research☆21Sep 12, 2025Updated 10 months ago
- 个人学习的医疗大模型微调项目☆37Dec 3, 2025Updated 8 months ago
- My implementation in TianChi CCKS 2025 pdf QA multimodal competition☆19Aug 27, 2025Updated 11 months ago
- 零基础入门推荐系统 - 新闻推荐 Top2☆44Mar 19, 2025Updated last year
- Open Ended Medical Reinforcement Learning☆65Mar 15, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Python 实现 MCP client / service☆81May 7, 2025Updated last year
- LLM大模型(重点)以及搜广推等 AI 算法中手写的面试题,(非 LeetCode),比如 Self-Attention, AUC等,一般比 LeetCode 更考察一个人的综合能力,又更贴近业务和基础知识一点☆600May 4, 2026Updated 3 months ago
- Awesome List for Agentic RL☆1,765Jul 23, 2026Updated 2 weeks ago
- 三元三小时手敲大模型☆558Mar 12, 2026Updated 4 months ago
- When Reasoning Meets Its Laws☆38Jan 2, 2026Updated 7 months ago
- ☆41Nov 16, 2025Updated 8 months ago
- 超简单使用监督微调SFT和强化学习RL去训练领域Agent☆34Oct 20, 2025Updated 9 months ago
- LLM中相关RLHF算法实现与学习☆15Apr 13, 2025Updated last year
- Reproducing and studying RL algorithms for LLM agents, including GRPO, GSPO, DAPO, OPD, Search-R1, ReTool, ALFWorld and beyond.☆134Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Nov 29, 2025Updated 8 months ago
- ☆33Sep 4, 2025Updated 11 months ago
- 🎯Awesome-AgenticRAG_DeepResearch: A curated list of resources on Agentic RAG & DeepResearch. 学习参考关于AgenticRAG、DeepResearch的发展相关论文☆33Jul 23, 2026Updated 2 weeks ago
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆64Apr 13, 2025Updated last year
- 杭高院自然语言处理课程2023☆26Nov 22, 2023Updated 2 years ago
- Local-first interview recording review reports with a Codex skill and CLI.☆77May 16, 2026Updated 2 months ago
- 科大讯飞多模态RAG图文问答挑战赛☆74Aug 4, 2025Updated last year
- [CVPR 2025] PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models☆60Jan 30, 2026Updated 6 months ago
- 一个基于 **LangChain 1.0**、**阿里通义 Qwen** 和 **MCP(amap-maps)** 的旅游规划 Agent Demo,支持: - 并行调用高德地图相关工具(搜索 POI、路线规划、天气查询) - 中间件级别的城市约束、工具预算控制、输出城市…☆40Jun 26, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Huggingface PPO Demo☆30Sep 7, 2025Updated 11 months ago
- Lightning-responsive CosyVoice streaming API based on FastAPI.☆28Apr 27, 2026Updated 3 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆22,877Updated this week
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆14,862Jun 14, 2026Updated last month
- 武汉大学国家网安院软件安全☆16Dec 9, 2024Updated last year
- 机器学习笔记本 Mechine Learning notebook☆25Apr 24, 2025Updated last year
- ☆11Jul 14, 2024Updated 2 years ago