Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)
☆23Mar 1, 2026Updated 6 months ago
Alternatives and similar repositories for Learning-to-Draft
Users that are interested in Learning-to-Draft are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated 2 years ago
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆56Jul 15, 2025Updated last year
- WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario (COLING 2025)☆13Jan 5, 2025Updated last year
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Empowering Small VLMs to Think with Dynamic Memorization and Exploration☆19Mar 18, 2026Updated 5 months ago
- MemOCR: an OCR-driven visual memory agent.☆34May 17, 2026Updated 3 months ago
- [ICML 2026] Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decoding☆27Aug 9, 2026Updated last month
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- this is a work about UpliftRec☆10Dec 10, 2024Updated last year
- 🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs☆35Aug 14, 2026Updated last month
- [ICML 2025] AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism☆20Jul 14, 2025Updated last year
- Tiny-R2: A hybrid architecture integrating SWA, CSA, HCA, mHC, and DSMoE under the DeepSeek V4 design paradigm, enabling single-GPU OPD p…☆49May 30, 2026Updated 3 months ago
- PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)☆34Jun 10, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The official implementation of the paper "MotifRetro: Exploring the Combinability-Consistency Trade-offs in retrosynthesis via Dynamic Mo…☆11Jun 25, 2023Updated 3 years ago
- [ICML 2026 Spotlight] UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models☆29Updated this week
- [ICLR 2026] CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing☆17Jan 31, 2026Updated 7 months ago
- Development repository for the Digital Terraria Lab implementation of the Sugarscape agent-based societal simulation.☆21Sep 1, 2026Updated last week
- Code for "Multi-level Relevance Document Identifier Learning for Generative Retrieval". ACL 2025.☆24Nov 5, 2025Updated 10 months ago
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts☆18Apr 24, 2026Updated 4 months ago
- Source code of the paper "OpSparse: a Highly Optimized Framework for Sparse General Matrix Multiplication on GPUs"☆16Aug 23, 2022Updated 4 years ago
- Accurate and fast KV cache compression with a gating mechanism☆29Jul 27, 2026Updated last month
- ☆25Dec 11, 2021Updated 4 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Diffusion-based generative drug-like molecular editing with chemical natural language☆18Dec 22, 2024Updated last year
- C++ package to store Matrix Market (.mtx) file format sparse matrices in Compressed Row Storage (CSR) format.☆17Oct 16, 2019Updated 6 years ago
- Hierarchical Speculative Decoding is the SOTA verification algorithm for lossless accelerated LLM inference.☆25Apr 14, 2026Updated 5 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 3 months ago
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆49Mar 1, 2026Updated 6 months ago
- 计算机网络课程设计, 基于TCP协议的简易聊天机器人, 开发语言Python3, 初期版本只能在终端中运行(CLI), 最终完成版为客户端编写了"简陋"的图形界面, 使用Qt5(即PyQt5)实现☆10Jun 17, 2019Updated 7 years ago
- ☆43Jun 8, 2026Updated 3 months ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆30Apr 17, 2026Updated 4 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆46Jul 2, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICML 2025] CommVQ: Commutative Vector Quantization for KV Cache Compression☆28Sep 2, 2025Updated last year
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆22Mar 2, 2025Updated last year
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆119Dec 2, 2025Updated 9 months ago
- ☆13Jul 22, 2024Updated 2 years ago
- ☆67Jul 3, 2026Updated 2 months ago
- ☆21Sep 5, 2024Updated 2 years ago
- pruning vision models in torch☆17Dec 5, 2025Updated 9 months ago