☆108Jul 24, 2025Updated last year
Alternatives and similar repositories for study_rlhf
Users that are interested in study_rlhf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆173Mar 18, 2026Updated 6 months ago
- 复现大模型相关算法及一些学习记录☆3,533Jul 2, 2026Updated 2 months ago
- 面向个人投资者金融投资研究与辅助决策agent系统☆28Jul 9, 2026Updated 2 months ago
- llm & rl☆294Oct 24, 2025Updated 10 months ago
- ☆44Nov 22, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 个人学习的医疗大模型微调项目☆38Dec 3, 2025Updated 9 months ago
- Beyond Basic RAG, Empowering Real-Time Deep Research☆22Sep 12, 2025Updated last year
- 零基础入门推荐系统 - 新闻推荐 Top2☆44Mar 19, 2025Updated last year
- LLM大模型(重点)以及搜广推等 AI 算法中手写的面试题,(非 LeetCode),比如 Self-Attention, AUC等,一般比 LeetCode 更考察一个人的综合能力,又更贴近业务和基础知识一点☆614May 4, 2026Updated 4 months ago
- Python 实现 MCP client / service☆82May 7, 2025Updated last year
- Minimal reproduction of OneRec☆1,831Sep 7, 2026Updated last week
- ☆60Dec 23, 2025Updated 8 months ago
- Awesome List for Agentic RL☆1,844Updated this week
- 三元三小时手敲大模型☆611Mar 12, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- When Reasoning Meets Its Laws☆38Jan 2, 2026Updated 8 months ago
- ☆41Nov 16, 2025Updated 10 months ago
- 超简单使用监督微调SFT和强化学习RL去训练领域Agent☆35Oct 20, 2025Updated 11 months ago
- ☆15Nov 29, 2025Updated 9 months ago
- 武汉大学国家网络安全学院2021级操作系统期末大实验☆12Jan 2, 2024Updated 2 years ago
- ☆33Sep 4, 2025Updated last year
- 🎯Awesome-AgenticRAG_DeepResearch: A curated list of resources on Agentic RAG & DeepResearch. 学习参考关于AgenticRAG、DeepResearch的发展相关论文☆33Jul 23, 2026Updated last month
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆63Apr 13, 2025Updated last year
- Local-first interview recording review reports with a Codex skill and CLI.☆81May 16, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 杭高院自然语言处理课程2023☆26Nov 22, 2023Updated 2 years ago
- 科大讯飞多模态RAG图文问答挑战赛☆75Aug 4, 2025Updated last year
- Huggingface PPO Demo☆31Sep 7, 2025Updated last year
- ☆24Jun 12, 2026Updated 3 months ago
- Lightning-responsive CosyVoice streaming API based on FastAPI.☆28Aug 17, 2026Updated last month
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,492Updated this week
- Reproducing and studying RL algorithms for LLM agents, including GRPO, GSPO, DAPO, OPD, Search-R1, ReTool, ALFWorld and beyond.☆378Updated this week
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆15,133Jun 14, 2026Updated 3 months ago
- 武汉大学国家网安院软件安全☆16Dec 9, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 一个很小很小的RAG系统☆385Apr 29, 2025Updated last year
- 从零复现 minimind👉minimind-v☆401Dec 24, 2025Updated 8 months ago
- ☆17Apr 15, 2026Updated 5 months ago
- ☆31Feb 27, 2026Updated 6 months ago
- Transferring Genshin PVs into a freehand style with Diffusion Model.☆10Jun 5, 2024Updated 2 years ago
- ☆13Aug 9, 2023Updated 3 years ago
- ☆10Apr 15, 2023Updated 3 years ago