这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。
☆63Apr 13, 2025Updated last year
Alternatives and similar repositories for open-r1-reprod
Users that are interested in open-r1-reprod are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen2.5 0.5B GRPO☆87Feb 16, 2025Updated last year
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆876Feb 18, 2025Updated last year
- 从零预训练LLM、SFT、RLHF、DPO笔记整理+面试问题☆21Sep 2, 2024Updated 2 years ago
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use☆30Mar 13, 2026Updated 6 months ago
- 本项目旨在利用LangChain和大语言模型(如ZhipuAI)开发一个智能数据库问答系统。 该系统能够通过自然语言理解用户的查询请求,自动生成相应的SQL语句并执行,最后将查询结果以自然语言 形式返回用户。☆15Jul 31, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 🏆🏆 「大模型」All in one & All from scratch. 🌍🌍 收集、清洗数据,训练Tokenizer,预训练、SFT、GRPO!☆56Aug 12, 2025Updated last year
- Automatically exported from code.google.com/p/transducersaurus☆11Apr 1, 2015Updated 11 years ago
- This repository is associated with the research paper titled ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large…☆15Jun 4, 2025Updated last year
- Train a tiny LLaMA model from scratch to repeat your words using Reinforcement Learning from Human Feedback (RLHF)☆18May 23, 2024Updated 2 years ago
- code for EMNLP 2022 paper Better Few-Shot Relation Extraction with Label Prompt Dropout☆26Nov 8, 2024Updated last year
- 日期时间实体识别☆11Sep 10, 2020Updated 6 years ago
- ☆16Oct 28, 2025Updated 11 months ago
- ☆11Mar 5, 2024Updated 2 years ago
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆54Jul 15, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆44Mar 6, 2025Updated last year
- 收集整理大模型面试题☆13Aug 29, 2024Updated 2 years ago
- Code space for L4DC paper "State-wise Safe Reinforcement Learning With Pixel Observations"☆11Apr 5, 2024Updated 2 years ago
- Code for ACM MM 2024 paper "A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning"☆19Dec 5, 2024Updated last year
- 目标:构建一个更符合语言学的小而美的 llama 分词器,支持中英日三国语言☆19Jun 2, 2024Updated 2 years ago
- Code for L4DC 2022 paper: Joint Synthesis of Safety Certificate and Safe Control Policy Using Constrained Reinforcement Learning.☆14Jul 31, 2023Updated 3 years ago
- 超简单复现Deepseek-R1-Zero和Deepseek-R1,以「24点游戏」为例。通过zero-RL、SFT以及SFT+RL,以激发LLM的自主验证反思能力。 About Clean, minimal, accessible reproduction of Dee…☆35Apr 5, 2025Updated last year
- Code for paper Towards Mitigating LLM Hallucination via Self Reflection☆30Oct 9, 2023Updated 3 years ago
- A curated list of resources on on-policy distillation☆25Apr 13, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICLR 2026] Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs☆47May 20, 2025Updated last year
- Documentation at☆14Mar 27, 2025Updated last year
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making☆17Oct 23, 2025Updated 11 months ago
- AlphaGo Zero Reinforcement Learning Sokoban Solver☆11Jun 20, 2018Updated 8 years ago
- LLM 101: 一起入门大语言模型 课程网站☆15Feb 2, 2025Updated last year
- ☆15Aug 7, 2025Updated last year
- ☆109Jul 24, 2025Updated last year
- [ISSTA 2025] A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing☆13Feb 9, 2025Updated last year
- ☆11Sep 9, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation☆30Sep 23, 2026Updated 2 weeks ago
- A collection of publications that works on code models but beyond focusing on the accuracies.☆12Jun 30, 2023Updated 3 years ago
- Official repo for Directional Self-supervised Learning for Heavy Image Augmentations [CVPR2022]☆12Jun 29, 2022Updated 4 years ago
- Code for "Positional Diffusion: Ordering Unordered Sets with Diffusion Probabilistic Models"☆19Mar 21, 2023Updated 3 years ago
- 基于DPO算法微调语言大模型,简单好上手。☆53Jul 3, 2024Updated 2 years ago
- ☆13Jul 5, 2024Updated 2 years ago
- Evaluate state-of-the-art sparse embedding models on the LIMIT dataset (`limit-small` and `limit`) from google's paper `On the Theoretica…☆16Sep 4, 2025Updated last year