[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
☆54Jun 17, 2025Updated last year
Alternatives and similar repositories for DeFT
Users that are interested in DeFT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32Mar 24, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆68Updated this week
- Implement some method of LLM KV Cache Sparsity☆41Jun 6, 2024Updated 2 years ago
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆117Dec 2, 2025Updated 8 months ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Python Script to Open SJTU Dormitory Smart Lock☆10Sep 12, 2022Updated 3 years ago
- Legacy Code of ZJU Campus App for iOS☆11Jan 31, 2024Updated 2 years ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- ☆12Sep 4, 2021Updated 4 years ago
- ☆20Dec 24, 2024Updated last year
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆214Mar 18, 2026Updated 4 months ago
- 面向多平台编译优化的深度学习中间表示☆10Oct 28, 2024Updated last year
- [EMNLP 2025] DiagramEval: Evaluating LLM-Generated Diagrams via Graphs☆17Nov 1, 2025Updated 9 months ago
- ☆37Aug 7, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Disaggregated serving system for Large Language Models (LLMs).☆828Apr 6, 2025Updated last year
- ☆19Aug 10, 2024Updated 2 years ago
- Collect papers related to personalized text generation☆18Sep 6, 2021Updated 4 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆281Aug 31, 2024Updated last year
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆176Feb 27, 2026Updated 5 months ago
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation☆255Dec 16, 2024Updated last year
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆403Apr 22, 2025Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆84Jul 14, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for Aesop: Paraphrase Generation with Adaptive Syntactic Control (EMNLP 2021)☆26Jan 17, 2022Updated 4 years ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆117Aug 4, 2026Updated last week
- [NeurIPS 2024] Activating Self-Attention for Multi-Scene Absolute Pose Regression☆14Feb 24, 2025Updated last year
- Stateful LLM Serving☆105Mar 11, 2025Updated last year
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆149Nov 10, 2025Updated 9 months ago
- STREAMer: Benchmarking remote volatile and non-volatile memory bandwidth☆18Aug 21, 2023Updated 2 years ago
- A low-latency & high-throughput serving engine for LLMs☆516Jan 8, 2026Updated 7 months ago
- An innovative method expediting LLMs via streamlined semi-autoregressive generation and draft verification.☆29Apr 15, 2025Updated last year
- ☆46Oct 15, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆17May 10, 2024Updated 2 years ago
- Course website for Systems Verification Fall 2024☆14Jul 10, 2025Updated last year
- Distributed MoE in a Single Kernel [NeurIPS '25]☆281May 5, 2026Updated 3 months ago
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆55Jul 15, 2025Updated last year
- LLM Evaluation Framework for Hardware Design Using Python-Embedded DSLs☆18Aug 26, 2024Updated last year
- [ACL 25] The Low-cost Long Context Understanding Benchmark for Large Language Models (Outstanding Paper Award)☆23Jul 30, 2025Updated last year
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length☆171Dec 23, 2025Updated 7 months ago