zzbright1998 / SentenceKVView on GitHub
Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache compression method that organizes cache at sentence level using semantic similarity.
15Sep 29, 2025Updated 10 months ago

Alternatives and similar repositories for SentenceKV

Users that are interested in SentenceKV are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?