[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
☆54Jun 17, 2025Updated last year
Alternatives and similar repositories for DeFT
Users that are interested in DeFT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32Mar 24, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆68Jun 24, 2026Updated 3 weeks ago
- ☆63May 19, 2025Updated last year
- Implement some method of LLM KV Cache Sparsity☆41Jun 6, 2024Updated 2 years ago
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆115Dec 2, 2025Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- Python Script to Open SJTU Dormitory Smart Lock☆10Sep 12, 2022Updated 3 years ago
- Legacy Code of ZJU Campus App for iOS☆11Jan 31, 2024Updated 2 years ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated 11 months ago
- ☆12Sep 4, 2021Updated 4 years ago
- ☆20Dec 24, 2024Updated last year
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆210Mar 18, 2026Updated 4 months ago
- 面向多平台编译优化的深度学习中间表示☆10Oct 28, 2024Updated last year
- A recommendation model kernel optimizing system☆12Jun 5, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [EMNLP 2025] DiagramEval: Evaluating LLM-Generated Diagrams via Graphs☆17Nov 1, 2025Updated 8 months ago
- ☆37Aug 7, 2025Updated 11 months ago
- Disaggregated serving system for Large Language Models (LLMs).☆826Apr 6, 2025Updated last year
- Collect papers related to personalized text generation☆18Sep 6, 2021Updated 4 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆281Aug 31, 2024Updated last year
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆174Feb 27, 2026Updated 4 months ago
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation☆255Dec 16, 2024Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆84Jul 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for Aesop: Paraphrase Generation with Adaptive Syntactic Control (EMNLP 2021)☆26Jan 17, 2022Updated 4 years ago
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆71Mar 17, 2026Updated 4 months ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆115Jun 28, 2025Updated last year
- Stateful LLM Serving☆105Mar 11, 2025Updated last year
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆146Nov 10, 2025Updated 8 months ago
- ☆15Jan 7, 2022Updated 4 years ago
- STREAMer: Benchmarking remote volatile and non-volatile memory bandwidth☆18Aug 21, 2023Updated 2 years ago
- A low-latency & high-throughput serving engine for LLMs☆512Jan 8, 2026Updated 6 months ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆121Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆401Apr 22, 2025Updated last year
- ☆45Oct 15, 2025Updated 9 months ago
- ☆17May 10, 2024Updated 2 years ago
- Course website for Systems Verification Fall 2024☆14Jul 10, 2025Updated last year
- Distributed MoE in a Single Kernel [NeurIPS '25]☆272May 5, 2026Updated 2 months ago
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆55Jul 15, 2025Updated last year
- LLM Evaluation Framework for Hardware Design Using Python-Embedded DSLs☆18Aug 26, 2024Updated last year