这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。
☆62Apr 13, 2025Updated last year
Alternatives and similar repositories for open-r1-reprod
Users that are interested in open-r1-reprod are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Sep 18, 2025Updated 8 months ago
- ☆11Feb 23, 2023Updated 3 years ago
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆55Jul 15, 2025Updated 10 months ago
- [ICLR 2026] Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs☆46May 20, 2025Updated last year
- ☆13May 16, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Train a tiny LLaMA model from scratch to repeat your words using Reinforcement Learning from Human Feedback (RLHF)☆18May 23, 2024Updated 2 years ago
- The first large scale formally verified reasoning dataset for Verilog☆21May 16, 2025Updated last year
- ☆106Jul 24, 2025Updated 10 months ago
- Code for ACM MM 2024 paper "A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning"☆19Dec 5, 2024Updated last year
- This repository is associated with the research paper titled ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large…☆15Jun 4, 2025Updated last year
- [IROS 2024 Oral Pitch] PyTorch Implementation of "Dual-Branch Graph Transformer Network for 3D Human Mesh Reconstruction from Video"☆15Jul 19, 2024Updated last year
- Code space for L4DC paper "State-wise Safe Reinforcement Learning With Pixel Observations"☆11Apr 5, 2024Updated 2 years ago
- Documentation at☆14Mar 27, 2025Updated last year
- ☆15Mar 30, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆19Sep 7, 2025Updated 9 months ago
- AlphaGo Zero Reinforcement Learning Sokoban Solver☆11Jun 20, 2018Updated 7 years ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆68Jan 28, 2026Updated 4 months ago
- [NAACL 2025] LLM-Supported Natural Language to Bash Translation☆16Apr 28, 2026Updated last month
- A simple utility for doing RISC-V HPM perf monitoring.☆18May 8, 2017Updated 9 years ago
- 数据管理平台(DataMan)是完全免费且开源的,任何人都可以无限制的修改代码以及部署服务,这对于很多想要对数据管理的应用平台来说是一个很好的选择:低廉的成本换回的是高效的管理方案,同时又有健康的生态提供支持。☆13Feb 25, 2022Updated 4 years ago
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model (ICCV 2025)☆37Sep 4, 2025Updated 9 months ago
- Code and data for COLING2024 paper "Characteristic AI Agents via Large Language Models".☆25Nov 29, 2024Updated last year
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- ☆16May 6, 2024Updated 2 years ago
- MPLS VPNs (VPLS, VPWS, L3VPN) on eNSP using Huawei Routers☆11Feb 11, 2020Updated 6 years ago
- 集成Qwen与DeepSeek等先进大语言模型,支持纯LLM+分类层模式及LLM+LoRA+分类层模式,使用transformers模块化设计和训练便于根据需要调整或替换组件。☆21Sep 1, 2025Updated 9 months ago
- [NeurIPS 2025 Spotlight] Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning☆86Sep 19, 2025Updated 8 months ago
- Code for "Positional Diffusion: Ordering Unordered Sets with Diffusion Probabilistic Models"☆18Mar 21, 2023Updated 3 years ago
- ☆28Sep 25, 2024Updated last year
- AI 应用示例合集☆114Jun 3, 2024Updated 2 years ago
- ☆24Dec 21, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆34Oct 16, 2021Updated 4 years ago
- Official repository for gathering data of Revisit Human-Scene Interaction via Space Occupancy (ECCV 2024).☆30Sep 29, 2024Updated last year
- ☆64May 4, 2025Updated last year
- SCoRe: Training Language Models to Self-Correct via Reinforcement Learning☆16May 14, 2026Updated 3 weeks ago
- Fine-tuned LLMs generate accurate 3D human avatars from textual descriptions using the SMPL-X model, enhancing customization and simulati…☆38Feb 5, 2025Updated last year
- Broodwar replays dumper (using BWAPI)☆13Apr 19, 2012Updated 14 years ago
- The open-source code for the NeurIPS 2025 paper, "Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learn…☆54Jan 5, 2026Updated 5 months ago