Accurate and fast KV cache compression with a gating mechanism
☆28Jul 27, 2026Updated last week
Alternatives and similar repositories for FastKVzip
Users that are interested in FastKVzip are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 5 months ago
- Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents☆26Feb 1, 2026Updated 6 months ago
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆45Mar 10, 2025Updated last year
- [ECCV 2024] Official PyTorch implementation of "HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts"☆20Nov 22, 2024Updated last year
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICML 2025] Official PyTorch implementation of "NegMerge: Sign-Consensual Weight Merging for Machine Unlearning"☆16Nov 25, 2025Updated 8 months ago
- Official PyTorch implementation of "Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier Data" (NeurIPS'23)☆15Dec 4, 2023Updated 2 years ago
- ☆15Apr 25, 2025Updated last year
- LLM KV cache compression made easy☆1,153Updated this week
- [TMLR'25] Official implementation for "Large-Scale Targeted Cause Discovery via Learning from Simulated Data"☆28Sep 30, 2025Updated 10 months ago
- A sparse-first inference engine (sparsevllm). It also contains DeltaKV compressor training + evaluation tooling (deltakv).☆63Updated this week
- Design and analyze optimal deep learning models.