☆49Mar 15, 2025Updated last year
Alternatives and similar repositories for Block-Attention
Users that are interested in Block-Attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆48Oct 16, 2025Updated 11 months ago
- Recomputation-free, context-independent KV caching for large language models☆37Aug 12, 2026Updated last month
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- ☆26Jan 16, 2025Updated last year
- Code repo for "CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs".☆17Sep 15, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The repo for our paper: Enhancing LLM-Based Agents via Global Planning and Hierarchical Execution (NCIIP 2025 Best Paper)☆17Aug 18, 2025Updated last year
- ☆15Feb 10, 2026Updated 7 months ago
- ☆205Jul 15, 2025Updated last year
- The official repository of paper "Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework".☆18Sep 22, 2025Updated last year
- Code and Data for "Language Modeling with Editable External Knowledge"☆39Jun 19, 2024Updated 2 years ago
- This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.☆14Apr 9, 2024Updated 2 years ago
- [SIGMOD 2025] PQCache: Product Quantization-based KVCache for Long Context LLM Inference