Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding
☆18May 6, 2026Updated 3 months ago
Alternatives and similar repositories for ReST-RL
Users that are interested in ReST-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A library of replicated state machine algorithms is based on Viewstamped Replication Revisited☆14Feb 6, 2021Updated 5 years ago
- ☆10May 18, 2023Updated 3 years ago
- ☆34Jun 30, 2026Updated 2 months ago
- SciGLM: Training Scientific Language Models with Self-Reflective Instruction Annotation and Tuning (NeurIPS D&B Track 2024)☆89Feb 25, 2024Updated 2 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆30Mar 30, 2026Updated 5 months ago
- A toy implementation of the dependently typed lambda calculus known as λΠ☆12Jan 29, 2020Updated 6 years ago
- A high availability distributed filesystem built on FoundationDB and fuse.☆22Mar 13, 2023Updated 3 years ago
- Download AudioSet for Vision-Audio-Text Pre-training☆13May 16, 2022Updated 4 years ago
- Sharing work on resumption monad☆12Sep 18, 2012Updated 13 years ago
- A Collection of Public Recommender System Dataset☆11Feb 10, 2021Updated 5 years ago
- Hood debugger, based on the idea of observing functions and structures as they are evaluated.☆20Jun 3, 2018Updated 8 years ago
- code repo for paper accepted in ICML 2023☆13Oct 19, 2023Updated 2 years ago
- Official code for "Self-supervised learning with rotation-invariant kernels"☆14Jun 5, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Blackbox Fuzzing of Distributed Systems with Multi-Dimensional Inputs and Symmetry-Based Feedback Pruning☆13Mar 7, 2025Updated last year
- A simulator for the Paxos Protocol for consensus in distributed systems☆21Dec 19, 2012Updated 13 years ago
- 根据Qwen2(Qwen1.5)模型生成qwen2 MoE模型的工具☆15Mar 29, 2024Updated 2 years ago
- ☆15Nov 14, 2023Updated 2 years ago
- Simple, Incremental SAT Solving as a Haskell Library☆15Aug 31, 2016Updated 9 years ago
- Manipulating Common Intermediate Language AST in Haskell☆22Nov 12, 2016Updated 9 years ago
- [ICML 2025] LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models☆16Nov 4, 2025Updated 9 months ago
- ☆19Jun 6, 2014Updated 12 years ago
- PELA: Learning Parameter-Efficient Models with Low-Rank Approximation [CVPR 2024]☆19Apr 14, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search in ACL'25☆105Jun 16, 2025Updated last year
- Interfaces in ziglang☆24Jan 8, 2023Updated 3 years ago
- Tiny LLM chatbot powered by Qwen2.5:0.5B-Instruct. It can run locally on a laptop. Depoloyment on Google Cloud is also supported.☆20Apr 1, 2025Updated last year
- Tools for working with the TREC CAR dataset.☆38Jul 12, 2025Updated last year
- A simple environment for writing and experimenting with hand-written CUDA PTX kernels.☆18Sep 11, 2025Updated 11 months ago
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 2 months ago
- UnitEval is a benchmarking and evaluation tools for AutoDev Coder.☆14Jan 2, 2024Updated 2 years ago
- ZetaSQL - Analyzer Framework for SQL☆14Oct 23, 2024Updated last year
- Convert TLA+ output (and values) into JSON☆28Mar 3, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for "CREAM: Consistency Regularized Self-Rewarding Language Models", ICLR 2025.☆30Feb 17, 2025Updated last year
- Tools for conformance monitoring on Kubernetes with TLA+☆23Jun 26, 2024Updated 2 years ago
- ☆23Jul 10, 2025Updated last year
- Parallel Token Prediction for Language Models (ICLR 2026)☆30Apr 27, 2026Updated 4 months ago
- Model-based clustering package for mixed data☆13May 21, 2026Updated 3 months ago
- Code and data for paper "(How) do Language Models Track State?"☆27Mar 31, 2025Updated last year
- Modern utility library and typescript typings for building JSON Schema documents☆14Nov 28, 2025Updated 9 months ago