这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。
☆63Apr 13, 2025Updated last year
Alternatives and similar repositories for open-r1-reprod
Users that are interested in open-r1-reprod are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆876Feb 18, 2025Updated last year
- 集成Qwen与DeepSeek等先进大语言模型,支持纯LLM+分类层模式及LLM+LoRA+分类层模式,使用transformers模块化设计和训练便于根据需要调整或替换组件。☆21Sep 1, 2025Updated last year
- 从零预训练LLM、SFT、RLHF、DPO笔记整理+面试问题☆21Sep 2, 2024Updated 2 years ago
- Load balancing based on reinforcement learning.☆11Oct 11, 2020Updated 5 years ago
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use☆30Mar 13, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 本项目旨在利用LangChain和大语言模型(如ZhipuAI)开发一个智能数据库问答系统。 该系统能够通过自然语言理解用户的查询请求,自动生成相应的SQL语句并执行,最后将查询结果以自然语言 形式返回用户。☆15Jul 31, 2024Updated 2 years ago
- ☆12Jan 5, 2023Updated 3 years ago
- Based on NS-3, design new GPSR routing protocols.☆13May 27, 2018Updated 8 years ago
- 放弃幻想、时刻准备、随时面试☆14Dec 17, 2025Updated 9 months ago
- Run TRex with PPO☆38May 17, 2025Updated last year
- kaggle:otto competition☆24Feb 13, 2023Updated 3 years ago
- Open-source repository for the OOPSLA'24 paper "CYCLE: Learning to Self-Refine Code Generation"☆10Mar 8, 2024Updated 2 years ago
- Build a simple basic multimodal large model from scratch. 从零搭建一个简单的基础多模态大模型🤖☆49Jun 19, 2024Updated 2 years ago
- This repository is associated with the research paper titled ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large…☆15Jun 4, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Github Repository for "Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees". (NeurIPS 2024)☆11Nov 30, 2025Updated 9 months ago
- Train a tiny LLaMA model from scratch to repeat your words using Reinforcement Learning from Human Feedback (RLHF)☆18May 23, 2024Updated 2 years ago
- 日期时间实体识别☆11Sep 10, 2020Updated 6 years ago
- ☆20Sep 29, 2024Updated last year
- Solution to kaggle competition OTTO – Multi-Objective Recommender System: https://www.kaggle.com/competitions/otto-recommender-system☆23Feb 2, 2023Updated 3 years ago
- 清华大学电子系--大一下小学期python大作业--一个很简陋的基于的机器学习的人脸识别系统☆10Sep 2, 2022Updated 4 years ago
- ☆15Mar 30, 2025Updated last year
- ☆14Apr 4, 2025Updated last year
- ☆11Mar 5, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- some notes for opensource llm technical reorts☆20Mar 6, 2026Updated 6 months ago
- ☆44Mar 6, 2025Updated last year
- Code space for L4DC paper "State-wise Safe Reinforcement Learning With Pixel Observations"☆11Apr 5, 2024Updated 2 years ago
- Code for ACM MM 2024 paper "A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning"☆19Dec 5, 2024Updated last year
- A ROS 2 package providing a collection of interfaces for hydrodynamic parameters.☆17Aug 13, 2026Updated last month
- 目标:构建一个更符合语言学的小而美的 llama 分词器,支持中英日三国语言☆19Jun 2, 2024Updated 2 years ago
- 2020Tianchi Competition News Recommendation☆11Jan 26, 2021Updated 5 years ago
- A curated list of resources on on-policy distillation☆25Apr 13, 2026Updated 5 months ago
- Greedy Perimeter Stateless Routing (GPSR) implement on NS3 platform☆21Aug 17, 2017Updated 9 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Packet Routing Simulator for Multi-Agent Reinforcement Learning☆27Jul 23, 2024Updated 2 years ago
- [ICLR 2026] Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs☆47May 20, 2025Updated last year
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- This is a repository used by individuals to experiment and reproduce the pre-training process of LLM.☆504May 1, 2025Updated last year
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making☆17Oct 23, 2025Updated 10 months ago
- ☆108Jul 24, 2025Updated last year
- CodeBERT based mutation testing tool.☆13Nov 10, 2025Updated 10 months ago