[ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.
☆56May 2, 2025Updated last year
Alternatives and similar repositories for RSD
Users that are interested in RSD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆75Oct 2, 2025Updated 10 months ago
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 2 months ago
- Make reasoning models scalable☆51Jun 2, 2026Updated 2 months ago
- Self-Hinting Language Models Enhance Reinforcement Learning☆28Mar 28, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- PyTorch implementation of StableMask (ICML'24)☆15Jun 27, 2024Updated 2 years ago
- [ACL 2026 Main] Official Repo for Paper "Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Ali…☆16Jul 1, 2026Updated last month
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- ☆18Apr 23, 2025Updated last year
- Continuous Pipelined Speculative Decoding☆22May 25, 2026Updated 2 months ago
- ☆20Mar 18, 2026Updated 5 months ago
- A Collection for Distributed Reinforcement Learning Papers☆18Sep 24, 2025Updated 10 months ago
- [NeurIPS 2025] Official PyTorch implementation for the paper AutoJudge: Judge Decoding Without Manual Annotation☆21Dec 22, 2025Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official PyTorch implementation for Training Unbiased Diffusion Models From Biased Dataset in ICLR 2024.☆15Apr 30, 2024Updated 2 years ago
- Open Source Projects from Pallas Lab☆21Oct 10, 2021Updated 4 years ago
- This is the official repo for Towards Uncertainty-Aware Language Agent.☆31Aug 15, 2024Updated 2 years ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆32Feb 17, 2025Updated last year
- ☆63May 19, 2025Updated last year
- ☆15Apr 26, 2025Updated last year
- VidKV: Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models☆25Mar 26, 2025Updated last year
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMs☆227Nov 30, 2025Updated 8 months ago
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 4 months ago
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.☆21Jan 24, 2025Updated last year
- ☆276May 14, 2025Updated last year
- This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or re…☆43Sep 22, 2024Updated last year
- [FPGA'26 Best Paper Nomination] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving☆36Aug 7, 2026Updated 2 weeks ago
- [EMNLP 2024 Findings🔥] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context In…☆103Nov 9, 2024Updated last year
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination.☆22Jul 18, 2025Updated last year
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural Networks☆11Oct 21, 2022Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code for "RSQ: Learning from Important Tokens Leads to Better Quantized LLMs"☆23Mar 25, 2026Updated 4 months ago
- ☆25Oct 31, 2024Updated last year
- ☆14Nov 2, 2025Updated 9 months ago
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆17Jul 18, 2024Updated 2 years ago
- [CVPR 2025] Pytorch implementation of the paper "Learning to Highlight Audio by Watching Movies"☆15Oct 1, 2025Updated 10 months ago
- Paper list for Efficient Reasoning.☆901May 29, 2026Updated 2 months ago
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length☆171Dec 23, 2025Updated 7 months ago