Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)
☆22Mar 1, 2026Updated 5 months ago
Alternatives and similar repositories for Learning-to-Draft
Users that are interested in Learning-to-Draft are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated last year
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆55Jul 15, 2025Updated last year
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆127Jul 25, 2026Updated last week
- [ICLR 2026] Empowering Small VLMs to Think with Dynamic Memorization and Exploration☆18Mar 18, 2026Updated 4 months ago
- MemOCR: an OCR-driven visual memory agent.☆33May 17, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆84Jul 14, 2025Updated last year
- [ICML 2025] AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism☆20Jul 14, 2025Updated last year
- The official implementation of the paper "MotifRetro: Exploring the Combinability-Consistency Trade-offs in retrosynthesis via Dynamic Mo…☆11Jun 25, 2023Updated 3 years ago
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 8 months ago
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts☆16Apr 24, 2026Updated 3 months ago
- [ICML 2026 Spotlight] UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models☆27May 1, 2026Updated 3 months ago
- [ICLR 2026] CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing☆16Jan 31, 2026Updated 6 months ago
- Code for "Multi-level Relevance Document Identifier Learning for Generative Retrieval". ACL 2025.☆24Nov 5, 2025Updated 9 months ago
- Source code of the paper "OpSparse: a Highly Optimized Framework for Sparse General Matrix Multiplication on GPUs"☆16Aug 23, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Accurate and fast KV cache compression with a gating mechanism☆28Jul 27, 2026Updated last week
- Language modeling based on Penn Treebank (RNN/LSTM, Pytorch)☆16Dec 11, 2019Updated 6 years ago
- ☆25Dec 11, 2021Updated 4 years ago
- ☆25Apr 8, 2026Updated 3 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆34May 26, 2026Updated 2 months ago
- The first continuous diffusion language model that rivals discrete counterparts on standard language modeling benchmarks like LM1B and Op…☆87Jun 14, 2026Updated last month
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆42Mar 1, 2026Updated 5 months ago
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse☆131Mar 14, 2026Updated 4 months ago
- 计算机网络课程设计, 基于TCP协议的简易聊天机器人, 开发语言Python3, 初期版本只能在终端中运行(CLI), 最终完成版为客户端编写了"简陋"的图形界面, 使用Qt5(即PyQt5)实现☆10Jun 17, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆42Jun 8, 2026Updated last month
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆28Apr 17, 2026Updated 3 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆43Jul 2, 2026Updated last month
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆22Mar 2, 2025Updated last year
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆117Dec 2, 2025Updated 8 months ago
- ☆13Jul 22, 2024Updated 2 years ago
- Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought☆20Feb 10, 2026Updated 5 months ago
- ☆21Sep 5, 2024Updated last year
- [CVPR 2023] Better “CMOS” Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-Resolution☆11Mar 19, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Sequential Diffusion Language Model (SDLM) enhances pre-trained autoregressive language models by adaptively determining generation lengt…☆98Dec 27, 2025Updated 7 months ago
- ☆25Mar 14, 2026Updated 4 months ago
- Nvidia TensorRT implementation of AdderNet for edge deployment☆10Nov 19, 2020Updated 5 years ago
- Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference☆56Mar 28, 2026Updated 4 months ago
- OpenH264 decode raw h264 demo.☆10Jul 8, 2017Updated 9 years ago
- (ECCV2026) Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs☆15Updated this week
- ☆16Jun 17, 2026Updated last month