A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.
☆170Mar 28, 2026Updated 5 months ago
Alternatives and similar repositories for grpo_reproduce
Users that are interested in grpo_reproduce are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆41Nov 16, 2025Updated 9 months ago
- Pretrain、Posttrain、RAG、Agent等大模型相关的基础项目合集☆39Dec 7, 2025Updated 8 months ago
- ☆18Mar 15, 2026Updated 5 months ago
- Local-first interview recording review reports with a Codex skill and CLI.☆79May 16, 2026Updated 3 months ago
- Fine-tune Qwen2.5-VL-7B on custom visual QA tasks using LoRA + Accelerate, supporting single/multi-GPU training on COCO 2014 dataset.☆30Apr 28, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。☆5,760Jun 3, 2026Updated 2 months ago
- some notes for opensource llm technical reorts