zzbright1998 / SentenceKVView on GitHub
Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache compression method that organizes cache at sentence level using semantic similarity.
15Sep 29, 2025Updated 9 months ago

Alternatives and similar repositories for SentenceKV

Users that are interested in SentenceKV are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?