[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
☆55Jun 17, 2025Updated last year
Alternatives and similar repositories for DeFT
Users that are interested in DeFT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32Mar 24, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆68Aug 22, 2026Updated 3 weeks ago
- ☆68May 19, 2025Updated last year
- Implement some method of LLM KV Cache Sparsity☆41Jun 6, 2024Updated 2 years ago
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆120Dec 2, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆54May 12, 2026Updated 4 months ago
- Python Script to Open SJTU Dormitory Smart Lock☆10Sep 12, 2022Updated 4 years ago
- Legacy Code of ZJU Campus App for iOS☆11Jan 31, 2024Updated 2 years ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- ☆12Sep 4, 2021Updated 5 years ago
- ☆20Dec 24, 2024Updated last year
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆221Mar 18, 2026Updated 6 months ago
- 面向多平台编译优化的深度学习中间表示☆10Oct 28, 2024Updated last year
- A recommendation model kernel optimizing system☆14Jun 5, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [EMNLP 2025] DiagramEval: Evaluating LLM-Generated Diagrams via Graphs☆17Nov 1, 2025Updated 10 months ago
- ☆37Aug 7, 2025Updated last year
- ☆19Aug 10, 2024Updated 2 years ago
- Disaggregated serving system for Large Language Models (LLMs).☆838Apr 6, 2025Updated last year
- Collect papers related to personalized text generation☆18Sep 6, 2021Updated 5 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆281Aug 31, 2024Updated 2 years ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆179Feb 27, 2026Updated 6 months ago
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation☆255Dec 16, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆413Apr 22, 2025Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- Code for Aesop: Paraphrase Generation with Adaptive Syntactic Control (EMNLP 2021)☆26Jan 17, 2022Updated 4 years ago
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆77Mar 17, 2026Updated 6 months ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆118Aug 4, 2026Updated last month
- [NeurIPS 2024] Activating Self-Attention for Multi-Scene Absolute Pose Regression☆15Feb 24, 2025Updated last year
- Stateful LLM Serving☆106Mar 11, 2025Updated last year
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆152Nov 10, 2025Updated 10 months ago
- STREAMer: Benchmarking remote volatile and non-volatile memory bandwidth☆18Aug 21, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A low-latency & high-throughput serving engine for LLMs☆527Jan 8, 2026Updated 8 months ago
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- An innovative method expediting LLMs via streamlined semi-autoregressive generation and draft verification.☆29Apr 15, 2025Updated last year
- ☆46Oct 15, 2025Updated 11 months ago
- ☆17May 10, 2024Updated 2 years ago
- Course website for Systems Verification Fall 2024☆14Jul 10, 2025Updated last year
- Distributed MoE in a Single Kernel [NeurIPS '25]☆291May 5, 2026Updated 4 months ago