Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding
☆18May 6, 2026Updated 3 months ago
Alternatives and similar repositories for ReST-RL
Users that are interested in ReST-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DataSciBench: An LLM Agent Benchmark for Data Science (Findings of ACL 2026)☆66Jan 21, 2026Updated 6 months ago
- A library of replicated state machine algorithms is based on Viewstamped Replication Revisited☆14Feb 6, 2021Updated 5 years ago
- cpp write language detect model☆11Sep 22, 2021Updated 4 years ago
- ☆10May 18, 2023Updated 3 years ago
- What makes Paxos tick?☆11Apr 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SciGLM: Training Scientific Language Models with Self-Reflective Instruction Annotation and Tuning (NeurIPS D&B Track 2024)☆88Feb 25, 2024Updated 2 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- First neural GPT aligned with text and speech. Welcome to join us to make better foundation model in neural modality.☆14Oct 30, 2024Updated last year
- ☆16Sep 25, 2025Updated 10 months ago
- A high availability distributed filesystem built on FoundationDB and fuse.☆22Mar 13, 2023Updated 3 years ago
- QuickCheck extension for higher-order properties☆19Feb 14, 2022Updated 4 years ago
- A Collection of Public Recommender System Dataset☆11Feb 10, 2021Updated 5 years ago
- ☆10Nov 1, 2024Updated last year
- Ice is a rapid information extraction customizer☆15Apr 26, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of ResFit, Residual Off-Policy RL for Finetuning Behavior Cloning Policies☆17Sep 29, 2025Updated 10 months ago
- code repo for paper accepted in ICML 2023☆13Oct 19, 2023Updated 2 years ago
- Official code for "Self-supervised learning with rotation-invariant kernels"☆14Jun 5, 2023Updated 3 years ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- ☆16Jul 15, 2025Updated last year
- A simulator for the Paxos Protocol for consensus in distributed systems☆21Dec 19, 2012Updated 13 years ago
- ☆15Nov 14, 2023Updated 2 years ago
- Run TLC in cmd☆15Jan 20, 2026Updated 6 months ago
- Simple, Incremental SAT Solving as a Haskell Library☆15Aug 31, 2016Updated 9 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- modified cutlass☆16Oct 26, 2020Updated 5 years ago
- [ICML 2025] LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models☆16Nov 4, 2025Updated 9 months ago
- 中文 电商 电脑 手机 相机 槽填充 数据集☆13Jan 14, 2020Updated 6 years ago
- Tiny LLM chatbot powered by Qwen2.5:0.5B-Instruct. It can run locally on a laptop. Depoloyment on Google Cloud is also supported.☆20Apr 1, 2025Updated last year
- Some Basic Computing and Simulating Tool in Math Modeling☆12Nov 6, 2019Updated 6 years ago
- Executive Memory for Coherent Long-Horizon Reasoning!☆86Jan 14, 2026Updated 6 months ago
- Tools for working with the TREC CAR dataset.☆38Jul 12, 2025Updated last year
- [TOIS 2023] On the User Behavior Leakage from Recommender System Exposure☆19Nov 7, 2023Updated 2 years ago
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The code repo for ICASSP 2023 Paper "MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning"☆26May 18, 2023Updated 3 years ago
- [arXiv 2025] SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning☆70Dec 17, 2025Updated 7 months ago
- ☆23May 24, 2025Updated last year
- UnitEval is a benchmarking and evaluation tools for AutoDev Coder.☆14Jan 2, 2024Updated 2 years ago
- View Zookeeper znode tree in a browser☆25Nov 18, 2015Updated 10 years ago
- Haskell implementation of the Edinburgh Logical Framework☆34Jan 12, 2026Updated 6 months ago
- Code for "CREAM: Consistency Regularized Self-Rewarding Language Models", ICLR 2025.☆29Feb 17, 2025Updated last year