Fast inference from large lauguage models via speculative decoding
β917Aug 22, 2024Updated last year
Alternatives and similar repositories for LLMSpeculativeSampling
Users that are interested in LLMSpeculativeSampling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Explorations into some recent techniques surrounding speculative decodingβ304Dec 22, 2024Updated last year
- π° Must-read papers and blogs on Speculative Decoding β‘οΈβ1,246Jun 2, 2026Updated last week
- Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Headsβ2,748Jun 25, 2024Updated last year
- REST: Retrieval-Based Speculative Decoding, NAACL 2024β219Mar 5, 2026Updated 3 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)β397Apr 22, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**β228Feb 13, 2025Updated last year
- Multi-Candidate Speculative Decodingβ41Apr 22, 2024Updated 2 years ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).β2,386Feb 20, 2026Updated 3 months ago
- [ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decodingβ1,337Mar 6, 2025Updated last year
- Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Trainingβ1,888Updated this week
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decodingβ281Aug 31, 2024Updated last year
- [NeurIPS'23] Speculative Decoding with Big Little Decoderβ98Feb 6, 2024Updated 2 years ago
- [ICLR 2025] SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Accelerationβ68Feb 21, 2025Updated last year
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Lengthβ163Dec 23, 2025Updated 5 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- πA curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.πβ5,270Apr 20, 2026Updated last month
- FlashInfer: Kernel Library for LLM Servingβ5,760Updated this week
- β30May 24, 2025Updated last year
- [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.β518Aug 1, 2024Updated last year
- Simple implementation of Speculative Sampling in NumPy for GPT-2.β99Aug 20, 2023Updated 2 years ago
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Seβ¦β843Mar 6, 2025Updated last year
- [MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Accelerationβ3,556Jul 17, 2025Updated 10 months ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabiliβ¦β4,086Updated this week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.β5,523Updated this week
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β28Mar 14, 2024Updated 2 years ago
- β16Aug 19, 2024Updated last year
- β67Dec 3, 2024Updated last year
- USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inferenceβ672May 21, 2026Updated 3 weeks ago
- β318Jul 10, 2025Updated 11 months ago
- Transformer related optimization, including BERT, GPTβ6,421Mar 27, 2024Updated 2 years ago
- Disaggregated serving system for Large Language Models (LLMs).β819Apr 6, 2025Updated last year
- [ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Headsβ541Feb 10, 2025Updated last year
- β157Mar 4, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- scalable and robust tree-based speculative decoding algorithmβ378Jan 28, 2025Updated last year
- [MLSys'24] Atom: Low-bit Quantization for Efficient and Accurate LLM Servingβ341Jul 2, 2024Updated last year
- π° Must-read papers on KV Cache Compression (constantly updating π€).β713Apr 15, 2026Updated last month
- Implementation of the paper Fast Inference from Transformers via Speculative Decoding, Leviathan et al. 2023.β110Dec 2, 2024Updated last year
- β359Apr 2, 2024Updated 2 years ago
- [ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cacheβ405Nov 20, 2025Updated 6 months ago
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Modelsβ1,658Jul 12, 2024Updated last year