Fine-tunes a student LLM using teacher feedback for improved reasoning and answer quality. Implements GRPO with teacher-provided evaluations.
☆54May 7, 2025Updated last year
Alternatives and similar repositories for grpo-llm-evaluator
Users that are interested in grpo-llm-evaluator are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NanoGPT (124M) in 5 minutes☆16Feb 14, 2025Updated last year
- this is based on the paper Chain-of-Retrieval Augmented Generation☆15Mar 29, 2025Updated last year
- Official implementation of PolySkill, a framework that enables web agents to learn generalizable and compositional skills through polymor…☆16Jul 6, 2026Updated last month
- ☆19Mar 10, 2025Updated last year
- ☆15Jan 17, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Multilingual RAG benchmark.☆11Nov 22, 2024Updated last year
- ☆15Apr 26, 2025Updated last year
- Easy to deploy your LLM(large language model) server with no public address GPU machine.☆15Apr 30, 2024Updated 2 years ago
- Бенчмарк для оценки способности языковых моделей решать математические и физические задачи на русском языке☆22Nov 14, 2025Updated 9 months ago
- Optimizing Causal LMs through GRPO with weighted reward functions and automated hyperparameter tuning using Optuna☆60Oct 18, 2025Updated 10 months ago
- This repository helps you evaluate your models on the FreshStack benchmark!☆34Dec 9, 2025Updated 8 months ago
- An easy-to-understand framework for LLM samplers that rewind and revise generated tokens☆151Jan 7, 2026Updated 7 months ago
- Local LLM Agent with Guidance☆13May 26, 2023Updated 3 years ago
- [TMLR 2025] When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models☆126Mar 6, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Seamless Voice Interactions with LLMs☆12Oct 28, 2023Updated 2 years ago
- [NeurIPS 2024] Low rank memory efficient optimizer without SVD☆33Jul 1, 2025Updated last year
- NeuroBLAST v3 architecture code☆37Jan 6, 2026Updated 7 months ago
- OpenCoconut implements a latent reasoning paradigm where we generate thoughts before decoding.☆173Jan 16, 2025Updated last year
- Simulation of job offers and CVs with real-time processing, classification, and analytics using Kafka, Ray, Spark, and Databricks. Includ…☆14Dec 25, 2024Updated last year
- ☆20Aug 1, 2024Updated 2 years ago
- chatGPT 'Autonomous Agent' in Node.js, written/runs in Termux. Sandboxed REPL access, Termux:API interface, chain-of-thought Question-Obs…☆16May 12, 2023Updated 3 years ago
- Interpreting Learned Search and Planning: Reverse-engineering recurrent convolutional networks (DRC) that play Sokoban☆22Jun 29, 2025Updated last year
- A locally trained model of Stoney Nakoda has been developed and released. You can access the working model here or train your own instanc…☆10Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Your agent is powerful but it doesn't know you. VibeLens visualizes agent sessions, personalizes your agents, provides dashboard analytic…☆19Updated this week
- An introduction to LLM Sampling☆80Dec 15, 2024Updated last year
- Hercules: Attributable and Scalable Opinion Summarization (ACL 2023)☆20Nov 8, 2023Updated 2 years ago
- Modify Entropy Based Sampling to work with Mac Silicon via MLX☆49Nov 6, 2024Updated last year
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆17Feb 20, 2026Updated 5 months ago
- Condense source code for LLM analysis by extracting essential highlights, utilizing a simplified version of Paul Gauthier's repomap techn…☆14Mar 3, 2024Updated 2 years ago
- [WWW 2026 Oral] MoE-CL:Self-Evolving LLMs via Continual Instruction Tuning☆21Dec 1, 2025Updated 8 months ago
- ☆16Jun 25, 2024Updated 2 years ago
- Collect papers about Mamba (a selective state space model).☆15Aug 6, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆10Apr 25, 2024Updated 2 years ago
- ☆25Oct 10, 2025Updated 10 months ago
- [ACL 2025] Knowledge Unlearning for Large Language Models☆49Sep 18, 2025Updated 11 months ago
- Approximating the joint distribution of language models via MCTS☆22Nov 3, 2024Updated last year
- Exploring Applications of GRPO☆252Aug 25, 2025Updated 11 months ago
- minimal pytorch implementation of bm25 (with sparse tensors)☆105Oct 28, 2025Updated 9 months ago
- ☆16Mar 22, 2025Updated last year