LLM KV cache compression made easy
β1,142Jul 9, 2026Updated last week
Alternatives and similar repositories for kvpress
Users that are interested in kvpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π° Must-read papers on KV Cache Compression (constantly updating π€).β725Apr 15, 2026Updated 3 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3β4Γ reduction in memory and 2Γ decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)β224Feb 11, 2026Updated 5 months ago
- The Official Implementation of Ada-KV [NeurIPS 2025]β139Nov 26, 2025Updated 7 months ago
- β324Jul 10, 2025Updated last year
- FlashInfer: Kernel Library for LLM Servingβ5,988Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cacheβ418Nov 20, 2025Updated 8 months ago
- [NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attentionβ¦β1,221Apr 8, 2026Updated 3 months ago
- [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferenceβ400Jul 10, 2025Updated last year
- [ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inferenceβ310May 1, 2025Updated last year
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Seβ¦β850Mar 6, 2025Updated last year
- This repository serves as a comprehensive survey of LLM development, featuring numerous research papers along with their corresponding coβ¦β339Updated this week
- [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.β528Aug 1, 2024Updated last year
- Accurate and fast KV cache compression with a gating mechanismβ26Apr 5, 2026Updated 3 months ago
- [ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Headsβ539Feb 10, 2025Updated last year
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generationβ255Dec 16, 2024Updated last year
- Awesome-LLM-KV-Cache: A curated list of πAwesome LLM KV Cache Papers with Codes.β460Jun 17, 2026Updated last month
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layerβ10,732Updated this week
- KV cache compression for high-throughput LLM inferenceβ158Feb 5, 2025Updated last year
- π Efficient implementations for emerging model architecturesβ5,379Updated this week
- Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inferenceβ54Mar 28, 2026Updated 3 months ago
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β997Updated this week
- NVIDIA Inference Xfer Library (NIXL)β1,139Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ6,674Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,109Sep 4, 2024Updated last year
- [ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoringβ280Jul 6, 2025Updated last year
- Distributed Compiler based on Triton for Parallel Systemsβ1,494Updated this week
- Accelerating MoE with IO and Tile-aware Optimizationsβ732Jul 4, 2026Updated 2 weeks ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.β5,925Updated this week
- Efficient LLM Inference over Long Sequencesβ392Jun 25, 2025Updated last year
- Awesome LLM compression research papers and tools.β1,853Jun 30, 2026Updated 3 weeks ago
- Perplexity GPU Kernelsβ591Nov 7, 2025Updated 8 months ago
- β39Mar 17, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β47Nov 25, 2024Updated last year
- Code for paper: [ICLR2025 Oral] FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inferenceβ170Oct 13, 2025Updated 9 months ago
- Cold Compress is a hackable, lightweight, and open-source toolkit for creating and benchmarking cache compression methods built on top ofβ¦β153Aug 9, 2024Updated last year
- Unified KV Cache Compression Methods for Auto-Regressive Modelsβ1,352Jul 10, 2026Updated last week
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decodingβ281Aug 31, 2024Updated last year
- Fast low-bit matmul kernels in Tritonβ477Updated this week
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLMβ3,562Updated this week