Official implementation of "TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization" (Findings of ACL 2025).
☆21Jul 25, 2025Updated last year
Alternatives and similar repositories for TailorKV
Users that are interested in TailorKV are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆53May 13, 2024Updated 2 years ago
- ☆17Apr 17, 2025Updated last year
- Code for the note "NF4 Isn't Information Theoretically Optimal (and that's Good)☆20Jun 22, 2023Updated 3 years ago
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- [SIGMOD 2025] PQCache: Product Quantization-based KVCache for Long Context LLM Inference☆91Dec 7, 2025Updated 7 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)