Accurate and fast KV cache compression with a gating mechanism
☆29Jul 27, 2026Updated 3 weeks ago
Alternatives and similar repositories for FastKVzip
Users that are interested in FastKVzip are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆224Feb 11, 2026Updated 6 months ago
- Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents☆28Feb 1, 2026Updated 6 months ago
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆45Mar 10, 2025Updated last year
- [ECCV 2024] Official PyTorch implementation of "HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts"☆20Nov 22, 2024Updated last year
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICML 2025] Official PyTorch implementation of "NegMerge: Sign-Consensual Weight Merging for Machine Unlearning"☆16Nov 25, 2025Updated 8 months ago
- Official PyTorch implementation of "Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier Data" (NeurIPS'23)☆15Dec 4, 2023Updated 2 years ago
- ☆15Apr 25, 2025Updated last year
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆18Jul 26, 2026Updated 3 weeks ago
- A sparse-first inference engine (sparsevllm). It also contains DeltaKV compressor training + evaluation tooling (deltakv).☆69Updated this week
- ☆23Sep 24, 2025Updated 11 months ago
- [ECCV 2024] Official PyTorch implementation of LUT "Learning with Unmasked Tokens Drives Stronger Vision Learners"☆14Dec 1, 2024Updated last year
- Official PyTorch implementation of "GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance" (ICML 2025)☆55Apr 13, 2026Updated 4 months ago
- Pytorch implementation for "Compressed Context Memory For Online Language Model Interaction" (ICLR'24)☆63Apr 18, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆13Jun 4, 2024Updated 2 years ago
- ☆11May 24, 2024Updated 2 years ago
- ☆54Nov 5, 2024Updated last year
- ☆14Aug 30, 2023Updated 2 years ago
- [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models☆28Aug 11, 2026Updated last week
- ☆11Oct 9, 2019Updated 6 years ago
- Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Model…☆72Mar 7, 2024Updated 2 years ago
- Code for "FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge". EMNLP 2023.☆20Dec 25, 2023Updated 2 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆22Mar 1, 2026Updated 5 months ago
- ☆24Mar 7, 2025Updated last year
- WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling☆21Apr 11, 2026Updated 4 months ago
- ☆43Jun 8, 2026Updated 2 months ago
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models☆104Apr 20, 2026Updated 4 months ago
- ImageNet-12k subset of ImageNet-21k (fall11)☆23Jun 13, 2023Updated 3 years ago
- ☆23Jun 16, 2026Updated 2 months ago
- ☆14Jan 11, 2024Updated 2 years ago
- ☆22Sep 11, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for "DISCO: accurate Discrete Scale Convolutions"☆19Jul 22, 2022Updated 4 years ago
- ☆14Oct 3, 2024Updated last year
- Reinforcement Learning from Text Feedback☆47Feb 17, 2026Updated 6 months ago
- [ICCV'25] HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics☆37Sep 10, 2025Updated 11 months ago
- 📰 Must-read papers on KV Cache Compression (constantly updating 🤗).☆736Aug 16, 2026Updated last week
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆25Dec 5, 2025Updated 8 months ago
- ☆37Oct 10, 2024Updated last year