Fine-tunes a student LLM using teacher feedback for improved reasoning and answer quality. Implements GRPO with teacher-provided evaluations.
☆54May 7, 2025Updated last year
Alternatives and similar repositories for grpo-llm-evaluator
Users that are interested in grpo-llm-evaluator are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NanoGPT (124M) in 5 minutes☆16Feb 14, 2025Updated last year
- Official implementation of PolySkill, a framework that enables web agents to learn generalizable and compositional skills through polymor…☆18Jul 6, 2026Updated 2 months ago
- ☆19Mar 10, 2025Updated last year
- ☆17Feb 1, 2024Updated 2 years ago
- ☆15Apr 26, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆13Mar 25, 2026Updated 6 months ago
- Train and finutune text-to-speech models for Bengali and many other languages!☆18Apr 2, 2025Updated last year
- ☆17Sep 1, 2026Updated 3 weeks ago
- A simple one file python script that executes AI processes defined in YML.☆14Mar 26, 2023Updated 3 years ago
- A curated list of awesome sentiment analysis studies, in which attitude corresponds to the text position conveyed by Subject towards othe…☆19Mar 23, 2026Updated 6 months ago
- Optimizing Causal LMs through GRPO with weighted reward functions and automated hyperparameter tuning using Optuna☆60Oct 18, 2025Updated 11 months ago
- This repository helps you evaluate your models on the FreshStack benchmark!☆34Dec 9, 2025Updated 9 months ago
- An easy-to-understand framework for LLM samplers that rewind and revise generated tokens☆151Jan 7, 2026Updated 8 months ago
- Seamless Voice Interactions with LLMs☆12Oct 28, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [TMLR 2025] When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models☆126Mar 6, 2026Updated 6 months ago
- [NeurIPS 2024] Low rank memory efficient optimizer without SVD☆34Jul 1, 2025Updated last year
- Tensorflow tf.metrics tutorial☆12Aug 30, 2018Updated 8 years ago
- NeuroBLAST v3 architecture code☆37Jan 6, 2026Updated 8 months ago
- OpenCoconut implements a latent reasoning paradigm where we generate thoughts before decoding.☆173Jan 16, 2025Updated last year
- Generating Easy-to-Understand Referring Expressions for Target Identifications☆18Aug 30, 2019Updated 7 years ago
- ☆24Jan 22, 2025Updated last year
- ☆16Nov 16, 2025Updated 10 months ago
- Learning adapter weights from task descriptions☆21Nov 12, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆19Aug 1, 2024Updated 2 years ago
- ☆17Oct 11, 2023Updated 2 years ago
- Interpreting Learned Search and Planning: Reverse-engineering recurrent convolutional networks (DRC) that play Sokoban☆22Jun 29, 2025Updated last year
- An introduction to LLM Sampling☆80Dec 15, 2024Updated last year
- Hercules: Attributable and Scalable Opinion Summarization (ACL 2023)☆20Nov 8, 2023Updated 2 years ago
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆18Feb 20, 2026Updated 7 months ago
- All the content of my youtube channel : https://youtube.com/@florenzerstling?si=7t10PBr6MDha74PO☆14May 28, 2025Updated last year
- Condense source code for LLM analysis by extracting essential highlights, utilizing a simplified version of Paul Gauthier's repomap techn…☆14Mar 3, 2024Updated 2 years ago
- ☆23May 25, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Collect papers about Mamba (a selective state space model).☆15Aug 6, 2024Updated 2 years ago
- ☆11Apr 25, 2024Updated 2 years ago
- ☆25Oct 10, 2025Updated 11 months ago
- Exploring Applications of GRPO☆251Aug 25, 2025Updated last year
- minimal pytorch implementation of bm25 (with sparse tensors)☆105Oct 28, 2025Updated 11 months ago
- Implemention based on lightrag and nano-graphrag to connect with psql☆15Oct 28, 2024Updated last year
- Easy local FLUX.1 Inference☆11Aug 29, 2024Updated 2 years ago