[ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.
☆56May 2, 2025Updated last year
Alternatives and similar repositories for RSD
Users that are interested in RSD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20May 14, 2025Updated last year
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆75Oct 2, 2025Updated 11 months ago
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 3 months ago
- Make reasoning models scalable☆51Jun 2, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Self-Hinting Language Models Enhance Reinforcement Learning☆28Mar 28, 2026Updated 5 months ago
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- ☆34Oct 13, 2025Updated 10 months ago
- PyTorch implementation of StableMask (ICML'24)☆15Jun 27, 2024Updated 2 years ago
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- ☆18Apr 23, 2025Updated last year
- Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]☆19May 12, 2025Updated last year
- Continuous Pipelined Speculative Decoding☆22May 25, 2026Updated 3 months ago
- ☆20Mar 18, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A Collection for Distributed Reinforcement Learning Papers☆18Sep 24, 2025Updated 11 months ago
- [NeurIPS 2025] Official PyTorch implementation for the paper AutoJudge: Judge Decoding Without Manual Annotation☆21Dec 22, 2025Updated 8 months ago
- Official PyTorch implementation for Training Unbiased Diffusion Models From Biased Dataset in ICLR 2024.☆15Apr 30, 2024Updated 2 years ago
- Open Source Projects from Pallas Lab☆21Oct 10, 2021Updated 4 years ago
- This is the official repo for Towards Uncertainty-Aware Language Agent.☆31Aug 15, 2024Updated 2 years ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆32Feb 17, 2025Updated last year
- ☆15Apr 26, 2025Updated last year
- ☆68May 19, 2025Updated last year
- T5Patches is a set of tools for fast and targeted editing of generative language models built with T5X.☆12May 31, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts☆17Apr 24, 2026Updated 4 months ago
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMs☆226Nov 30, 2025Updated 9 months ago
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 4 months ago
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 5 months ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆55Oct 18, 2024Updated last year
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.☆21Jan 24, 2025Updated last year
- ☆275May 14, 2025Updated last year
- This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or re…☆43Sep 22, 2024Updated last year
- Implementation of "RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm"☆18Apr 11, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving☆38Aug 7, 2026Updated last month
- [EMNLP 2024 Findings🔥] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context In…☆103Nov 9, 2024Updated last year
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination.☆22Jul 18, 2025Updated last year
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural Networks☆11Oct 21, 2022Updated 3 years ago
- Code for "RSQ: Learning from Important Tokens Leads to Better Quantized LLMs"☆24Mar 25, 2026Updated 5 months ago
- ☆25Oct 31, 2024Updated last year