Repository for the COLM 2025 paper SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
☆19Jul 10, 2025Updated last year
Alternatives and similar repositories for SpecDec_pp
Users that are interested in SpecDec_pp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆50Oct 24, 2023Updated 2 years ago
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15May 19, 2026Updated 3 months ago
- Official Implementation of DART (DART: Low-Latency Parallel Drafting with Continuity-Aware Tree Pruning for Speculative Decoding, EMNLP26…☆70Aug 25, 2026Updated last week
- ☆31May 24, 2025Updated last year
- ☆19May 4, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆14May 9, 2024Updated 2 years ago
- [NeurIPS 2021] Source code for the paper "Qu-ANTI-zation: Exploiting Neural Network Quantization for Achieving Adversarial Outcomes"☆18Nov 9, 2021Updated 4 years ago
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆24Nov 21, 2024Updated last year
- Code for the EMNLP24 paper "A simple and effective L2 norm based method for KV Cache compression."☆19Dec 13, 2024Updated last year
- ☆11May 13, 2025Updated last year
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆56Jul 15, 2025Updated last year
- NUMA-Aware Reader-Writer Locks☆19Jun 12, 2014Updated 12 years ago
- Princeton University Ph.D. Dissertation Template☆20Apr 9, 2017Updated 9 years ago
- 自然语言处理大作业-三种中文分词方法的性能对比与评分☆22Jan 17, 2021Updated 5 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A universal workflow system for exactly-once DAGs☆23Jun 1, 2023Updated 3 years ago
- ☆49Sep 13, 2025Updated 11 months ago
- MACER: MAximizing CErtified Radius (ICLR 2020)☆31Jan 5, 2020Updated 6 years ago
- Check your grade automatically and send e-mail when new grade comes☆12Feb 7, 2018Updated 8 years ago
- ☆68Nov 4, 2024Updated last year
- The Official Implementation of Ada-KV [NeurIPS 2025]☆139Nov 26, 2025Updated 9 months ago
- A fast and scalable distributed lock service using programmable switches.☆21Jul 30, 2024Updated 2 years ago
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,292Jun 27, 2026Updated 2 months ago
- Source code of IPA, https://escholarship.org/uc/item/2p0805dq☆13Jun 27, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of the paper "Pretraining Language Models to Ponder in Continuous Space"☆27Jul 21, 2025Updated last year
- [CVPR'25] Conformal prediction for vision-language models. Enhancing VLMs deployment with reliability gurarantees.☆21Jun 7, 2025Updated last year
- Official Repo for "SplitQuant / LLM-PQ: Resource-Efficient LLM Offline Serving on Heterogeneous GPUs via Phase-Aware Model Partition and …☆39Aug 29, 2025Updated last year
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation☆256Dec 16, 2024Updated last year
- A method for training neural networks that are provably robust to adversarial attacks. [IJCAI 2019]☆10Sep 3, 2019Updated 7 years ago
- REST: Retrieval-Based Speculative Decoding, NAACL 2024☆221Mar 5, 2026Updated 5 months ago
- KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches. EMNLP Findings 2024☆90Feb 27, 2025Updated last year
- ☆18Feb 5, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- 项目的主仓库☆26Sep 11, 2022Updated 3 years ago
- This is the implementation for IEEE S&P 2022 paper "Model Orthogonalization: Class Distance Hardening in Neural Networks for Better Secur…☆11Aug 24, 2022Updated 4 years ago
- [ICLR 2025] On Evluating the Durability of Safegurads for Open-Weight LLMs☆13Jun 20, 2025Updated last year
- [NeurIPS 2025 D&B] BackdoorDM: A Comprehensive Benchmark for Backdoor Learning in Diffusion Model☆29Aug 20, 2026Updated 2 weeks ago
- State of the art Chinese Word Segmentation with Bi-LSTMs☆27Apr 1, 2020Updated 6 years ago
- An LLM leaderboard for stateful agents☆21Oct 16, 2025Updated 10 months ago