Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding
☆18May 6, 2026Updated 2 months ago
Alternatives and similar repositories for ReST-RL
Users that are interested in ReST-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DataSciBench: An LLM Agent Benchmark for Data Science (Findings of ACL 2026)☆64Jan 21, 2026Updated 6 months ago
- ☆10May 18, 2023Updated 3 years ago
- Distributed, Replicated, Protocol-generic Key-value Store in Async Rust for SMR Protocols Research☆18Jul 5, 2026Updated 2 weeks ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- First neural GPT aligned with text and speech. Welcome to join us to make better foundation model in neural modality.☆14Oct 30, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Collection of Public Recommender System Dataset☆11Feb 10, 2021Updated 5 years ago
- ☆10Nov 1, 2024Updated last year
- Implementation of ResFit, Residual Off-Policy RL for Finetuning Behavior Cloning Policies☆17Sep 29, 2025Updated 9 months ago
- code repo for paper accepted in ICML 2023☆13Oct 19, 2023Updated 2 years ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 4 months ago
- [TMM 2019] Official Implementation for Hierarchical User Intent Graph Network for Multimedia Recommendation☆11Apr 6, 2026Updated 3 months ago
- A simulator for the Paxos Protocol for consensus in distributed systems☆21Dec 19, 2012Updated 13 years ago
- 根据Qwen2(Qwen1.5)模型生成qwen2 MoE模型的工具☆15Mar 29, 2024Updated 2 years ago
- Run TLC in cmd☆15Jan 20, 2026Updated 6 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- modified cutlass☆16Oct 26, 2020Updated 5 years ago
- ☆16Sep 19, 2023Updated 2 years ago
- Auditing agents for fine-tuning safety☆21Oct 21, 2025Updated 9 months ago
- [ICML 2025] LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models☆17Nov 4, 2025Updated 8 months ago
- ☆22Jun 11, 2024Updated 2 years ago
- Tiny LLM chatbot powered by Qwen2.5:0.5B-Instruct. It can run locally on a laptop. Depoloyment on Google Cloud is also supported.☆20Apr 1, 2025Updated last year
- Executive Memory for Coherent Long-Horizon Reasoning!☆84Jan 14, 2026Updated 6 months ago
- Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing☆18Sep 24, 2022Updated 3 years ago
- [arXiv 2025] SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning☆70Dec 17, 2025Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- View Zookeeper znode tree in a browser☆25Nov 18, 2015Updated 10 years ago
- UnitEval is a benchmarking and evaluation tools for AutoDev Coder.☆14Jan 2, 2024Updated 2 years ago
- Discrete event simulation of a cluster of machines with .NET Core async/await☆34May 3, 2018Updated 8 years ago
- Code for "CREAM: Consistency Regularized Self-Rewarding Language Models", ICLR 2025.☆29Feb 17, 2025Updated last year
- ZooKeeper server on top of FoundationDB☆26Mar 31, 2021Updated 5 years ago
- ☆24Aug 20, 2025Updated 11 months ago
- Our Clone of Orca used for experimentation☆19Oct 15, 2024Updated last year
- ☆29Aug 21, 2025Updated 11 months ago
- Parallel Token Prediction for Language Models (ICLR 2026)☆30Apr 27, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆25Aug 2, 2024Updated last year
- ☆21Feb 18, 2022Updated 4 years ago
- Code and data for paper "(How) do Language Models Track State?"☆26Mar 31, 2025Updated last year
- Model-based clustering package for mixed data☆13May 21, 2026Updated last month
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer (EMNLP 2025)☆12Apr 18, 2025Updated last year
- a very fast parser for sparse matrix at libsvm format☆10Nov 13, 2017Updated 8 years ago
- ☆11Aug 20, 2025Updated 11 months ago