LLM KV cache compression made easy
β1,208Sep 17, 2026Updated this week
Alternatives and similar repositories for kvpress
Users that are interested in kvpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π° Must-read papers on KV Cache Compression (constantly updating π€).β742Aug 16, 2026Updated last month
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3β4Γ reduction in memory and 2Γ decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)β225Feb 11, 2026Updated 7 months ago
- The Official Implementation of Ada-KV [NeurIPS 2025]β140Nov 26, 2025Updated 9 months ago
- β333Jul 10, 2025Updated last year
- FlashInfer: Kernel Library for LLM Servingβ6,452Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cacheβ433Nov 20, 2025Updated 9 months ago
- [NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attentionβ¦β1,229Sep 10, 2026Updated last week
- Awesome-LLM-KV-Cache: A curated list of πAwesome LLM KV Cache Papers with Codes.β466Jun 17, 2026Updated 3 months ago
- [ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inferenceβ314May 1, 2025Updated last year
- [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferenceβ407Jul 10, 2025Updated last year
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Seβ¦β861Mar 6, 2025Updated last year
- This repository serves as a comprehensive survey of LLM development, featuring numerous research papers along with their corresponding coβ¦β358Jul 16, 2026Updated 2 months ago
- [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.β534Aug 1, 2024Updated 2 years ago
- Accurate and fast KV cache compression with a gating mechanismβ31Jul 27, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Headsβ540Feb 10, 2025Updated last year
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generationβ255Dec 16, 2024Updated last year
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layerβ11,867Updated this week
- KV cache compression for high-throughput LLM inferenceβ161Feb 5, 2025Updated last year
- π Efficient implementations for emerging model architecturesβ5,766Updated this week
- Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inferenceβ58Mar 28, 2026Updated 5 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).β2,535Feb 20, 2026Updated 7 months ago
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β1,178Updated this week
- NVIDIA Inference Xfer Library (NIXL)β1,263Updated this week
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,433Updated this week
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,149Sep 4, 2024Updated 2 years ago
- [ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoringβ282Jul 6, 2025Updated last year
- Distributed Compiler and Optimized Parallel Kernelsβ1,546Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,544Updated this week
- Accelerating MoE with IO and Tile-aware Optimizationsβ769Aug 29, 2026Updated 3 weeks ago
- Efficient LLM Inference over Long Sequencesβ392Jun 25, 2025Updated last year
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.β6,614Updated this week
- Awesome LLM compression research papers and tools.β1,875Aug 27, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Perplexity GPU Kernelsβ608Nov 7, 2025Updated 10 months ago
- β40Mar 17, 2025Updated last year
- β47Nov 25, 2024Updated last year
- Code for paper: [ICLR2025 Oral] FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inferenceβ172Oct 13, 2025Updated 11 months ago
- Tile primitives for speedy kernelsβ3,716Sep 12, 2026Updated last week
- Cold Compress is a hackable, lightweight, and open-source toolkit for creating and benchmarking cache compression methods built on top ofβ¦β153Aug 9, 2024Updated 2 years ago
- Unified KV Cache Compression Methods for Auto-Regressive Modelsβ1,381Aug 13, 2026Updated last month