Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton
☆53May 12, 2026Updated 3 months ago
Alternatives and similar repositories for SAM-Decoding
Users that are interested in SAM-Decoding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL2025 Oral🔥]Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling☆30Nov 11, 2025Updated 9 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆407Apr 22, 2025Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- ☆20Dec 24, 2024Updated last year
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆55Jul 15, 2025Updated last year
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length☆172Dec 23, 2025Updated 8 months ago
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 4 months ago
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆217Mar 18, 2026Updated 5 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆44Jul 2, 2026Updated last month
- REST: Retrieval-Based Speculative Decoding, NAACL 2024☆221Mar 5, 2026Updated 5 months ago
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,291Jun 27, 2026Updated 2 months ago
- [NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.☆70Jan 28, 2026Updated 7 months ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆176Feb 27, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference☆55Jun 17, 2025Updated last year
- Simple speculative decoding technique, integrated in vLLM and transformers☆615Aug 23, 2024Updated 2 years ago
- [ICLR 2025] SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration☆71Feb 21, 2025Updated last year
- ☆69Dec 3, 2024Updated last year
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆468Updated this week
- The official implementation of "LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation"☆22Apr 22, 2025Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,514Feb 20, 2026Updated 6 months ago
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated 2 years ago
- ☆11Feb 5, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,131Updated this week
- [NeurIPS 2024] Fast Best-of-N Decoding via Speculative Rejection☆56Oct 29, 2024Updated last year
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆29Jul 4, 2026Updated last month
- TKDE'20 paper, "CONNA: Addressing Name Disambiguation on the Fly".☆15May 31, 2021Updated 5 years ago
- Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]☆19May 12, 2025Updated last year
- ☆38Jul 21, 2025Updated last year
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆137Jul 25, 2026Updated last month
- Ouroboros: Speculative Decoding with Large Model Enhanced Drafting (EMNLP 2024 main)☆117Mar 20, 2025Updated last year
- Nano SGLang☆16Jul 21, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆40Jun 23, 2026Updated 2 months ago
- This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Act…☆19Oct 25, 2024Updated last year
- ☆15Apr 11, 2024Updated 2 years ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 5 months ago
- Dirigent: Lightweight Serverless Orchestration☆45Aug 26, 2025Updated last year
- Chain of Thoughts (CoT) is so hot! so long! We need short reasoning process!☆72Apr 1, 2025Updated last year
- [ICML 2025] Official implementation of the paper "SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling". …☆21Nov 17, 2025Updated 9 months ago