[ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.
☆56May 2, 2025Updated last year
Alternatives and similar repositories for RSD
Users that are interested in RSD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20May 14, 2025Updated last year
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆75Oct 2, 2025Updated 9 months ago
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated last month
- Make reasoning models scalable☆51Jun 2, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Self-Hinting Language Models Enhance Reinforcement Learning☆27Mar 28, 2026Updated 4 months ago
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆84Jul 14, 2025Updated last year
- ☆34Oct 13, 2025Updated 9 months ago
- PyTorch implementation of StableMask (ICML'24)☆15Jun 27, 2024Updated 2 years ago
- [ACL 2026 Main] Official Repo for Paper "Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Ali…☆17Jul 1, 2026Updated 3 weeks ago
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- ☆18Apr 23, 2025Updated last year
- Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]☆19May 12, 2025Updated last year
- Continuous Pipelined Speculative Decoding☆22May 25, 2026Updated 2 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆20Mar 18, 2026Updated 4 months ago
- A Collection for Distributed Reinforcement Learning Papers☆18Sep 24, 2025Updated 10 months ago
- [NeurIPS 2025] Official PyTorch implementation for the paper AutoJudge: Judge Decoding Without Manual Annotation☆21Dec 22, 2025Updated 7 months ago
- Open Source Projects from Pallas Lab☆21Oct 10, 2021Updated 4 years ago
- This is the official repo for Towards Uncertainty-Aware Language Agent.☆31Aug 15, 2024Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆31Feb 17, 2025Updated last year
- ☆63May 19, 2025Updated last year
- VidKV: Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models☆25Mar 26, 2025Updated last year
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts☆16Apr 24, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMs☆225Nov 30, 2025Updated 8 months ago
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 2 months ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆54Oct 18, 2024Updated last year
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.☆21Jan 24, 2025Updated last year
- ☆275May 14, 2025Updated last year
- This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or re…☆43Sep 22, 2024Updated last year
- Implementation of "RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm"☆17Apr 11, 2025Updated last year
- [FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving☆34Feb 23, 2026Updated 5 months ago
- [EMNLP 2024 Findings🔥] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context In…☆103Nov 9, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination.☆21Jul 18, 2025Updated last year
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural Networks☆11Oct 21, 2022Updated 3 years ago
- ☆25Oct 31, 2024Updated last year
- ☆14Nov 2, 2025Updated 8 months ago
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆15Jul 18, 2024Updated 2 years ago
- JMLR Cover Letter Template☆10Dec 15, 2021Updated 4 years ago