Implementation of Speculative Sampling as described in "Accelerating Large Language Model Decoding with Speculative Sampling" by Deepmind
☆112Feb 29, 2024Updated 2 years ago
Alternatives and similar repositories for Speculative-Sampling
Users that are interested in Speculative-Sampling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Explorations into some recent techniques surrounding speculative decoding☆307Dec 22, 2024Updated last year
- Simple implementation of Speculative Sampling in NumPy for GPT-2.☆98Aug 20, 2023Updated 3 years ago
- Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads☆2,771Jun 25, 2024Updated 2 years ago
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,292Jun 27, 2026Updated 2 months ago
- Keyformer proposes KV Cache reduction through key tokens identification and without the need for fine-tuning☆59Mar 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**☆230Feb 13, 2025Updated last year
- A Collection for Distributed Reinforcement Learning Papers☆18Sep 24, 2025Updated 11 months ago
- Parallel waveform generation with DiffusionGAN☆17Mar 26, 2022Updated 4 years ago
- Repo hosting codes and materials related to speeding LLMs' inference using token merging.☆37Oct 9, 2025Updated 10 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 11 months ago
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- ☆16Jun 4, 2024Updated 2 years ago
- Official implementation of DGP-based multi-speaker speech synthesis with PyTorch☆24Mar 23, 2021Updated 5 years ago
- ☆24Mar 15, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Dec 14, 2022Updated 3 years ago
- ☆44Jun 2, 2024Updated 2 years ago
- All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks☆17Apr 24, 2024Updated 2 years ago
- [ICML 2024] CLLMs: Consistency Large Language Models☆416Nov 16, 2024Updated last year
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length☆172Dec 23, 2025Updated 8 months ago
- ☆14Dec 16, 2020Updated 5 years ago
- The official code for Dropping Backward Propagation (DropBP)☆32Oct 29, 2024Updated last year
- G2pw's inference speed is accelerated by about 8-10 times. Change loop generated predictive data to only once and model loop prediction b…☆14Dec 30, 2023Updated 2 years ago
- [SpeechCom Journal] Learning and controlling the source-filter representation of speech with a variational autoencoder☆46Apr 18, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding☆1,342Mar 6, 2025Updated last year
- Source code of APNet2, a vocoder☆61Nov 23, 2023Updated 2 years ago
- Multi-Candidate Speculative Decoding☆41Apr 22, 2024Updated 2 years ago
- [MLSys'24] Atom: Low-bit Quantization for Efficient and Accurate LLM Serving☆347Jul 2, 2024Updated 2 years ago
- NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates [WIP]☆25Jul 5, 2022Updated 4 years ago
- Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training☆1,899Updated this week
- GEAR: An Efficient KV Cache Compression Recipefor Near-Lossless Generative Inference of LLM☆186Jul 12, 2024Updated 2 years ago
- Code for EMNLP 2021 paper: Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting☆17Nov 30, 2021Updated 4 years ago
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆15Mar 11, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆28Mar 14, 2024Updated 2 years ago
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitin…☆73Jun 26, 2024Updated 2 years ago
- ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation☆25Oct 2, 2020Updated 5 years ago
- PyTorch Implementation of NCSOFT's FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis☆74Aug 3, 2021Updated 5 years ago
- ☆14Feb 3, 2026Updated 7 months ago
- This is the official implementation of our multi-channel multi-speaker multi-spatial neural audio codec architecture.☆55Mar 17, 2025Updated last year
- Codes for our paper "Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation" (EMNLP 2023 Findings)☆47Dec 9, 2023Updated 2 years ago