A comrephensive collection of learning from rewards in the post-training and test-time scaling of LLMs, with a focus on both reward models and learning strategies across training, inference, and post-inference stages.
☆75Jun 13, 2025Updated last year
Alternatives and similar repositories for learning-from-rewards-llm-papers
Users that are interested in learning-from-rewards-llm-papers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16May 18, 2026Updated 4 months ago
- Codes and data for KDD 2024 Research Track paper "ProCom: A Few-shot Targeted Community Detection Algorithm"☆11Aug 15, 2024Updated 2 years ago
- ☆15Dec 2, 2022Updated 3 years ago
- Official repository for EMNLP'22 paper: Grape: Knowledge Graph Enhanced Passage Reader for Open-domain Question Answering☆23Oct 20, 2022Updated 3 years ago
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- JudgeLRM: Large Reasoning Models as a Judge☆42May 6, 2026Updated 4 months ago
- instruction-following benchmark for large reasoning models☆49Apr 19, 2026Updated 5 months ago
- Code & Data for our Paper "Alleviating Hallucinations of Large Language Models through Induced Hallucinations"☆71Feb 27, 2024Updated 2 years ago
- Code for "Improving Translation Faithfulness of Large Language Models via Augmenting Instructions"☆12Aug 26, 2023Updated 3 years ago
- [NeurIPS 2024] Train LLMs with diverse system messages reflecting individualized preferences to generalize to unseen system messages☆53Aug 10, 2025Updated last year
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning.☆25Oct 7, 2025Updated 11 months ago
- GSM-Plus: Data, Code, and Evaluation for Enhancing Robust Mathematical Reasoning in Math Word Problems.☆67Jul 8, 2024Updated 2 years ago
- About The corresponding code from our paper " Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning…☆13Jan 14, 2026Updated 8 months ago
- Code for our SIGIR 2021 short paper "Lighter and Better: Low-Rank Decomposed Self-Attention Networks for Next-Item Recommendation."☆15May 5, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15May 27, 2025Updated last year
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆19Oct 20, 2025Updated 10 months ago
- The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.☆453Jul 11, 2025Updated last year
- The code for paper: Decoupled Planning and Execution: A Hierarchical Reasoning Framework for Deep Search [SIGIR 2026]☆65Jul 4, 2025Updated last year
- A Survey of Reinforcement Learning for Large Reasoning Models☆2,490Updated this week
- ☆36May 24, 2025Updated last year
- Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER (ACL 2022)☆44Apr 7, 2022Updated 4 years ago
- A curated list of cutting-edge research papers and resources on Long Chain-of-Thought (CoT) Reasoning with Tools.☆46Dec 17, 2025Updated 9 months ago
- ☆15Apr 11, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Source code for NAACL 2022 paper: Relation-Specific Attentions over Entity Mentions for Enhanced Document-Level Relation Extraction☆17May 9, 2022Updated 4 years ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆74Jan 28, 2026Updated 7 months ago
- verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in…☆2,320Jun 9, 2026Updated 3 months ago
- [AAAI 2026] ReCode: Reinforced Code Knowledge Editing for API Updates☆25Jul 1, 2025Updated last year
- 揣摩研习社关注自然语言和信息检索前沿技术,解读热门科技论文,分享实用科研工具,挖掘人工智能冰山之下的学术和应用价值!☆37Nov 4, 2022Updated 3 years ago
- ☆364Jul 29, 2025Updated last year
- ☆16Nov 19, 2021Updated 4 years ago
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?☆20Mar 9, 2025Updated last year
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆232Nov 27, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code and data for EMNLP 2023 paper "RAPL: A Relation-Aware Prototype Learning Approach for Few-Shot Document-Level Relation Extraction"☆18Mar 6, 2024Updated 2 years ago
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆93Dec 8, 2025Updated 9 months ago
- The repo of "Improving Seq2Seq Grammatical Error Correction via Decoding Interventions"☆32Jan 22, 2024Updated 2 years ago
- ☆14Apr 24, 2024Updated 2 years ago
- Code repository for the paper "The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Le…☆14Jan 16, 2025Updated last year
- ☆17Jul 17, 2026Updated 2 months ago
- ☆15Jan 14, 2026Updated 8 months ago