🐭 A tiny single-file implementation of Group Relative Policy Optimization (GRPO) as introduced by the DeepSeekMath paper
☆43Jun 28, 2025Updated last year
Alternatives and similar repositories for microGRPO
Users that are interested in microGRPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A minimal, hackable Vision-Language Model built on Karpathy’s nanochat — add image understanding and multimodal chat for under $200 in co…☆24Jul 14, 2026Updated 3 weeks ago
- Minimal hackable GRPO implementation☆345Jan 31, 2025Updated last year
- ☆10Dec 19, 2019Updated 6 years ago
- TD-Regularized Actor-Critic Methods☆37Dec 26, 2019Updated 6 years ago
- ☆14May 9, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Reinforcement learning training framework for entity-gym environments.☆17Mar 18, 2024Updated 2 years ago
- Implementation of Proximal Policy Optimization (PPO) for continuous action space (`Pendulum-v1` from gym) using tensorflow2.x and pytorch…☆12Aug 8, 2022Updated 4 years ago
- Aioli: A unified optimization framework for language model data mixing☆33Jan 17, 2025Updated last year
- Creating an environment to quickly train a variety of Deep Reinforcement Learning algorithms on Street Fighter 2 using tournaments betwee…☆23Mar 25, 2023Updated 3 years ago
- A Deep-Reinforcement-Learning-Based Scheduler for FPGA HLS☆15Feb 27, 2021Updated 5 years ago
- A simple tutorial to add medical reasoning using GRPO☆21Feb 10, 2025Updated last year
- ☆16Jul 9, 2025Updated last year
- Highly scalable 2D JAX physics engine.☆68Apr 20, 2026Updated 3 months ago
- Code for ThriftyDAgger☆15Dec 29, 2021Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- run deepseek v3 on a single node. Drops unused experts from memory.☆16Jan 26, 2025Updated last year
- A fully modular framework for modeling and optimizing analog neural networks☆21Jan 19, 2026Updated 6 months ago
- Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.☆18Dec 19, 2024Updated last year
- ☆33Jun 24, 2024Updated 2 years ago
- Code for the paper "VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit Assignment"☆192May 25, 2025Updated last year
- Kakao Mobility MCP Server for directions and transit information☆11Sep 14, 2025Updated 10 months ago
- Top 10 Data Centers & AI Infrastructure Security Risks☆17Jul 17, 2026Updated 3 weeks ago
- Use deep learning to learn Koopman operator and LQR for optimal control☆18Sep 28, 2020Updated 5 years ago
- ☆28Jun 2, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆11Jan 12, 2015Updated 11 years ago
- ☆12Jan 21, 2025Updated last year
- 삼각형의 실전! Triton☆16Feb 15, 2024Updated 2 years ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- ☆19Oct 12, 2025Updated 9 months ago
- Time-ordered UUIDv4☆20Jun 10, 2024Updated 2 years ago
- ☆13Aug 4, 2022Updated 4 years ago
- Repository for ICLR'23 Long-tailed Learning Requires Feature Learning☆10Feb 22, 2023Updated 3 years ago
- Implemention based on lightrag and nano-graphrag to connect with psql☆15Oct 28, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Alias-free Bessel function synthesis☆12Dec 12, 2014Updated 11 years ago
- Swift port of Welly BBS Client☆14Jun 8, 2021Updated 5 years ago
- Atari-style POMDPs☆34Updated this week
- PySOM - The Simple Object Machine Smalltalk implemented in Python☆18Jun 7, 2026Updated 2 months ago
- 2018年春季工科创IV-E:智能小车机器人☆10May 10, 2018Updated 8 years ago
- Code repository for the paper "The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Le…☆14Jan 16, 2025Updated last year
- The Pair App is employed by the Agency of Learning for team management and communication.☆11Apr 13, 2024Updated 2 years ago