π A tiny single-file implementation of Group Relative Policy Optimization (GRPO) as introduced by the DeepSeekMath paper
β44Jun 28, 2025Updated last year
Alternatives and similar repositories for microGRPO
Users that are interested in microGRPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Recreating the minimal training methods of DeepSeek-R1 for small langauge models.β21Feb 10, 2025Updated last year
- A minimal, hackable Vision-Language Model built on Karpathyβs nanochat β add image understanding and multimodal chat for under $200 in coβ¦β26Sep 21, 2026Updated 2 weeks ago
- Minimal hackable GRPO implementationβ347Jan 31, 2025Updated last year
- We open-source our layout level fast EM simulation tool, EMSim, to the public.β16Feb 8, 2024Updated 2 years ago
- TD-Regularized Actor-Critic Methodsβ37Dec 26, 2019Updated 6 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Create Custom GYM Environment for SUMO and reinforcement learning agantβ15May 5, 2023Updated 3 years ago
- Reinforcement learning training framework for entity-gym environments.β17Mar 18, 2024Updated 2 years ago
- Aioli: A unified optimization framework for language model data mixingβ34Jan 17, 2025Updated last year
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]β34Jan 23, 2025Updated last year
- Creating an environment to quickly train a variety of Deep Reinforcement Learning algorithms on Street Fighter 2 using tournaments betweeβ¦β23Mar 25, 2023Updated 3 years ago
- A Deep-Reinforcement-Learning-Based Scheduler for FPGA HLSβ15Feb 27, 2021Updated 5 years ago
- β14Aug 15, 2024Updated 2 years ago
- A simple tutorial to add medical reasoning using GRPOβ21Feb 10, 2025Updated last year
- β16Jul 9, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ADAPTIVE RESONANCE THEORY. Gail A. Carpenter and Stephen Grossbergβ10Feb 10, 2015Updated 11 years ago
- Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.β18Dec 19, 2024Updated last year
- β33Jun 24, 2024Updated 2 years ago
- Code for the paper "VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit Assignment"β195May 25, 2025Updated last year
- Kakao Mobility MCP Server for directions and transit informationβ11Sep 14, 2025Updated last year
- Use deep learning to learn Koopman operator and LQR for optimal controlβ18Sep 28, 2020Updated 6 years ago
- β28Jun 2, 2026Updated 4 months ago
- β10Aug 27, 2019Updated 7 years ago
- β11Jan 12, 2015Updated 11 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β12Jan 21, 2025Updated last year
- μΌκ°νμ μ€μ ! Tritonβ16Feb 15, 2024Updated 2 years ago
- Tiny evaluation of leading LLMs on competitive programming problemsβ14Apr 10, 2026Updated 6 months ago
- Large language model of Medical AI, General Medical AI (GMAI)β17Jan 30, 2024Updated 2 years ago
- Code for the paper "Learning to Schedule Joint Radar-Communication with Deep Multi-Agent Reinforcement Learning" as published in the IEEEβ¦β27Aug 10, 2022Updated 4 years ago
- Implemention based on lightrag and nano-graphrag to connect with psqlβ15Oct 28, 2024Updated last year
- [WACV 2024] Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining, WACV 2024β13Jan 3, 2024Updated 2 years ago
- β19Oct 12, 2025Updated 11 months ago
- Weight-Averaged Sharpness-Aware Minimization (NeurIPS 2022)β28Jan 13, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Repository for ICLR'23 Long-tailed Learning Requires Feature Learningβ10Feb 22, 2023Updated 3 years ago
- Alias-free Bessel function synthesisβ12Sep 29, 2026Updated last week
- A quick way to get started with Transformer Lensβ15Dec 13, 2023Updated 2 years ago
- Swift port of Welly BBS Clientβ14Jun 8, 2021Updated 5 years ago
- β20Sep 3, 2025Updated last year
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.β14Mar 20, 2024Updated 2 years ago
- β18Dec 5, 2017Updated 8 years ago