A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.
☆171Mar 28, 2026Updated 6 months ago
Alternatives and similar repositories for grpo_reproduce
Users that are interested in grpo_reproduce are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆41Nov 16, 2025Updated 10 months ago
- A very simple GRPO implement for reproducing r1-like LLM thinking.☆1,714Nov 21, 2025Updated 10 months ago
- 基于Qwen2+SFT+DPO的医疗问答系统,项目中使用了自定义的 SFTTrainer/DPOTrainer/TRPOTrainer用于训练,其次,项目还调用各种知识库工具(neo4j, milvus, LDA, 等)进行自动化训练数据生成。另外,使用 vllm 用于推理…☆90Apr 29, 2026Updated 5 months ago
- A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using …☆83Jul 6, 2026Updated 3 months ago
- 复现大模型相关算法及一些学习记录☆3,547Jul 2, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- 这是一个从零开始构建的强化学习人类反馈(RLHF)学习代码库,实现了 PPO、GRPO、GSPO 以及相关的策略优化算法,并提供了清晰、可复现的训练流程。由于文档是由latex文件转译过来,如果md文件渲染异常,请用VScode的md插件打开☆89Dec 19, 2025Updated 9 months ago
- WebResearcher: An Iterative Deep-Research Agent,迭代式深度研究智能体☆50Feb 13, 2026Updated 7 months ago
- 超简单使用监督微调SFT和强化学习RL去训练领域Agent☆35Oct 20, 2025Updated 11 months ago
- Multi-Modal-AI-Orchestrator (Reset version),AI Full-modal Full-agent:Text → Image → Music → Lights → Video, Includes "Scenario Director,…☆102Nov 5, 2025Updated 11 months ago
- 大学Latex答辩模版,当前包含川大、哈工大、中科大。☆11Jul 22, 2024Updated 2 years ago
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆876Feb 18, 2025Updated last year
- 2024百度商业AI技术创新大赛赛道一:基于大模型的广告检索全国一等奖获奖方案☆19Feb 23, 2025Updated last year
- My Solution and Notes for the Stanford CS336: LLM from scratch☆266Mar 23, 2026Updated 6 months ago
- official implementation of "CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusi…☆19Sep 5, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementing DeepSeek R1's GRPO algorithm from scratch☆1,904Apr 18, 2025Updated last year
- 将SmolVLM2的视觉头与Qwen3-0.6B模型进行了拼接微调☆612Sep 8, 2025Updated last year
- 🧠 Train a 64M-parameter LLM from scratch in just 2h!☆63,435Sep 22, 2026Updated 2 weeks ago
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,488Nov 13, 2025Updated 10 months ago
- Minimal reproduction of OneRec☆1,872Sep 7, 2026Updated last month
- 主要记录大语言大 模型(LLMs) 算法(应用)工程师相关的知识及面试题☆15,211Jun 14, 2026Updated 3 months ago
- ☆109Jul 24, 2025Updated last year
- 📚 从零开始构建大模型☆34,258Aug 8, 2026Updated 2 months ago
- 包含了LLM的一些手撕代码,如强化学习。可以帮助从代码层面深入理解原理,以及有助于准备大模型面试可能出 现的手撕。后续会更新Transformer等更多手撕☆136Mar 15, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 面向长程购物 Agent 的可复现后训练☆462Updated this week
- 集成Qwen与DeepSeek等先进大语言模型,支持纯LLM+分类层模式及LLM+LoRA+分类层模式,使用transformers模块化设计和训练便于根据需要调整或替换组件。☆22Sep 1, 2025Updated last year
- 基于轻量级 Qwen2.5-0.5B 和 SigLIP 的视觉语言多模态模型实现,包含训练和 SFT 代码。分享训练和 SFT 相关代码,记录一下探索和学习的过程。欢迎一起交流讨论~☆22Aug 31, 2025Updated last year
- 🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...☆701Mar 28, 2026Updated 6 months ago
- ☆523Oct 16, 2025Updated 11 months ago
- 收集为大模型面试准备的手撕代码☆56Mar 15, 2026Updated 6 months ago
- The code to reproduce CVPR 2021 paper "Towards Robust Classification Model by Counterfactual and Invariant Data Generation"☆16Jul 29, 2021Updated 5 years ago
- ☆881Jul 26, 2026Updated 2 months ago
- [CVPR2024] Distribution-aware Knowledge Prototyping for Non-exemplar Lifelong Person Re-identification☆18Jul 3, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Python code to solve robust multi-mode resource constrained project scheduling problem using Benders' decomposition approach vs compact m…☆16Jun 23, 2022Updated 4 years ago
- 每个人都能看懂的大模型知识分享,LLMs春/秋招大模型面试前必看,让你和面试官侃侃而谈☆7,425Aug 28, 2026Updated last month
- A set of examples based on verl for end-to-end RL training recipes.☆338Sep 22, 2026Updated 2 weeks ago
- ☆32Feb 27, 2026Updated 7 months ago
- 2025腾讯广告算法大赛相关代码——官方baseline☆39Nov 4, 2025Updated 11 months ago
- 🎓从0开始训练一个大模型Minimind项目的超详细解析,包括但不限于用到的架构,算法,以及大模型面试经验☆1,792May 25, 2026Updated 4 months ago
- My implementation of Stanford CS336 assignments.☆245Mar 15, 2026Updated 6 months ago