☆108Jul 24, 2025Updated 11 months ago
Alternatives and similar repositories for study_rlhf
Users that are interested in study_rlhf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆169Mar 18, 2026Updated 4 months ago
- Enhanced Search-R1 Implementation: Improved Compatibility and Modern Framework Integration☆28Dec 8, 2025Updated 7 months ago
- 复现大模型相关算法及一些学习记录☆3,463Jul 2, 2026Updated 2 weeks ago
- WebResearcher: An Iterative Deep-Research Agent,迭代式深度研究智能体☆49Feb 13, 2026Updated 5 months ago
- 面向个人投资者金融投资研究与辅助决策agent系统☆28Jul 9, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,123Nov 13, 2025Updated 8 months ago
- 🏥 从零基础到面试通关:20节课彻底搞懂MedicalGPT医疗大模型训练全流程 | PT/SFT/LoRA/RLHF/DPO/GRPO | 100+面试高频考点☆185Apr 1, 2026Updated 3 months ago
- 2025腾讯生成式推荐广告算法大赛(深海大菠萝)☆17Aug 3, 2025Updated 11 months ago
- 个人学习的医疗大模型微调项目☆37Dec 3, 2025Updated 7 months ago
- ☆45Nov 22, 2025Updated 7 months ago
- Beyond Basic RAG, Empowering Real-Time Deep Research☆20Sep 12, 2025Updated 10 months ago
- 零基础入门推荐系统 - 新闻推荐 Top2☆44Mar 19, 2025Updated last year
- Open Ended Medical Reinforcement Learning☆63Mar 15, 2026Updated 4 months ago
- Python 实现 MCP client / service☆81May 7, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- LLM大模型(重点)以及搜广推等 AI 算法中手写的面试题,(非 LeetCode),比如 Self-Attention, AUC等,一般比 LeetCode 更考察一个人的综合能力,又更贴近业务和基础知识一点☆598May 4, 2026Updated 2 months ago
- Minimal reproduction of OneRec☆1,709May 14, 2026Updated 2 months ago
- 三元三小时手敲大模型☆539Mar 12, 2026Updated 4 months ago
- Awesome List for Agentic RL☆1,701Jun 20, 2026Updated last month
- When Reasoning Meets Its Laws☆38Jan 2, 2026Updated 6 months ago
- An intelligent customer support system powered by LangGraph and LangChain that uses Retrieval-Augmented Generation (RAG) to provide accur…☆20Jul 25, 2025Updated 11 months ago
- ☆40Nov 16, 2025Updated 8 months ago
- 超简单使用监督微调SFT和强化学习RL去训练领域Agent☆35Oct 20, 2025Updated 9 months ago
- LLM中相关RLHF算法实现与学习☆15Apr 13, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Reproducing and studying RL algorithms for LLM agents, including PPO, GRPO, GSPO, DAPO, OPD and beyond.☆33Updated this week
- ☆15Nov 29, 2025Updated 7 months ago
- 武汉大学国家网络安全学院2021级操作系统期末大实验☆12Jan 2, 2024Updated 2 years ago
- 基于movielens-25m数据集的生成式推荐项目☆53Aug 6, 2025Updated 11 months ago
- ☆32Sep 4, 2025Updated 10 months ago
- 🎯Awesome-AgenticRAG_DeepResearch: A curated list of resources on Agentic RAG & DeepResearch. 学习参考关于AgenticRAG、DeepResearch的发展相关论文☆32Jun 12, 2026Updated last month
- Local-first interview recording review reports with a Codex skill and CLI.☆75May 16, 2026Updated 2 months ago
- 科大讯飞多模态RAG图文问答挑战赛☆74Aug 4, 2025Updated 11 months ago
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆64Apr 13, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 杭高院自然语言处理课程2023☆26Nov 22, 2023Updated 2 years ago
- Implementation of "M3amba: Memory Mamba is All You Need for Whole Slide Image Classification". CVPR2025☆12Feb 27, 2025Updated last year
- Huggingface PPO Demo☆29Sep 7, 2025Updated 10 months ago
- Lightning-responsive CosyVoice streaming API based on FastAPI.☆28Apr 27, 2026Updated 2 months ago
- 《大模型白盒子构建指南》:一个全手搓的Tiny-Universe☆4,968Feb 12, 2026Updated 5 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆22,571Updated this week
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆14,733Jun 14, 2026Updated last month