π° Must-read papers and blogs on Speculative Decoding β‘οΈ
β1,288Jun 27, 2026Updated last month
Alternatives and similar repositories for SpeculativeDecodingPapers
Users that are interested in SpeculativeDecodingPapers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)β403Apr 22, 2025Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).β2,500Feb 20, 2026Updated 5 months ago
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Lengthβ171Dec 23, 2025Updated 7 months ago
- Fast inference from large lauguage models via speculative decodingβ923Aug 22, 2024Updated last year
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β1,058Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Explorations into some recent techniques surrounding speculative decodingβ307Dec 22, 2024Updated last year
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.β214Mar 18, 2026Updated 4 months ago
- Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**β230Feb 13, 2025Updated last year
- [ICLR 2025] SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Accelerationβ70Feb 21, 2025Updated last year
- Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Headsβ2,762Jun 25, 2024Updated 2 years ago
- π° Must-read papers on KV Cache Compression (constantly updating π€).β731Apr 15, 2026Updated 3 months ago
- [ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decodingβ1,340Mar 6, 2025Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verificationβ84Jul 14, 2025Updated last year
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decodingβ281Aug 31, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A curated list for Efficient Large Language Modelsβ2,032Jun 17, 2025Updated last year
- πA curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.πβ5,448Jul 26, 2026Updated 2 weeks ago
- REST: Retrieval-Based Speculative Decoding, NAACL 2024β220Mar 5, 2026Updated 5 months ago
- Codes for our paper "Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation" (EMNLP 2023 Findings)β47Dec 9, 2023Updated 2 years ago
- Ouroboros: Speculative Decoding with Large Model Enhanced Drafting (EMNLP 2024 main)β117Mar 20, 2025Updated last year
- Multi-Candidate Speculative Decodingβ41Apr 22, 2024Updated 2 years ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automatonβ52May 12, 2026Updated 2 months ago
- scalable and robust tree-based speculative decoding algorithmβ377Jan 28, 2025Updated last year
- Awesome LLM compression research papers and tools.β1,859Jun 30, 2026Updated last month
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Paper list for Efficient Reasoning.β900May 29, 2026Updated 2 months ago
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)β56Mar 14, 2025Updated last year
- β68Dec 3, 2024Updated last year
- FlashInfer: Kernel Library for LLM Servingβ6,133Updated this week
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitinβ¦β73Jun 26, 2024Updated 2 years ago
- Reading notes on Speculative Decoding papersβ43Jun 2, 2026Updated 2 months ago
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Seβ¦β852Mar 6, 2025Updated last year
- [ICLR2025] Breaking Throughput-Latency Trade-off for Long Sequences with Speculative Decodingβ156Dec 4, 2024Updated last year
- Awesome-LLM-KV-Cache: A curated list of πAwesome LLM KV Cache Papers with Codes.β463Jun 17, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Trainingβ1,898Updated this week
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafterβ176Feb 27, 2026Updated 5 months ago
- Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"β1,070May 30, 2026Updated 2 months ago
- [ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cacheβ425Nov 20, 2025Updated 8 months ago
- A throughput-oriented high-performance serving framework for LLMsβ974Mar 29, 2026Updated 4 months ago
- Official implementation of βDomino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decodingβ.β129Jul 25, 2026Updated 2 weeks ago
- My learning notes for ML SYS.β6,842Updated this week