Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).
☆25Jun 9, 2026Updated 3 months ago
Alternatives and similar repositories for IceCache
Users that are interested in IceCache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of NOSA (EMNLP 2026 main)☆19Updated this week
- Segmented Code Adjustment Quantization (SAQ)☆28Sep 22, 2025Updated last year
- ☆17Updated this week
- [NeurIPS 2026] DMax: Aggressive Parallel Decoding for dLLMs☆132Updated this week
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"☆17Apr 30, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Jul 15, 2024Updated 2 years ago
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆26Dec 5, 2025Updated 9 months ago
- ☆17Mar 20, 2026Updated 6 months ago
- GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent [ICML 2026]☆41Sep 16, 2026Updated 2 weeks ago
- [ICML 2026] Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decoding☆27Aug 9, 2026Updated last month
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆16Aug 16, 2026Updated last month
- Repository for "Training Language Models To Explain Their Own Computations"☆38Jul 7, 2026Updated 2 months ago
- Code for "Reversal Q-Learning (RQL)" for Flow RL from Prior Data☆38Jun 17, 2026Updated 3 months ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆12May 19, 2025Updated last year
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 3 years ago
- ☆13Jan 7, 2025Updated last year
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15Updated this week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 5 months ago
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 6 months ago
- Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)☆26May 29, 2026Updated 4 months ago
- ☆14Jan 22, 2026Updated 8 months ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICML 2026] An Evaluation Suite for Chain-of-Thought Controllability☆56Mar 10, 2026Updated 6 months ago
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆25Jul 17, 2026Updated 2 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 6 months ago
- ☆29Mar 18, 2026Updated 6 months ago
- [CVPR 2026 Highlight] PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning☆33Sep 18, 2026Updated 2 weeks ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- Official Code Repository for OmniRetrieval☆35Jun 1, 2026Updated 4 months ago
- ☆17Mar 24, 2025Updated last year
- Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use☆73Sep 17, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆10Apr 26, 2023Updated 3 years ago
- Public repository for CleANN, an efficient fully-dynamic approximate nearest neighbor search index☆18Sep 18, 2025Updated last year
- A cross-modal vector index with fast construction on heterogeneous CPU-GPU environment. Published on DaMoN@SIGMOD 2025.☆16Jul 16, 2025Updated last year
- Repository for SIEVE: Effective Filtered Vector Search with Collection of Indexes VLDB '25 submission☆17Jul 30, 2025Updated last year
- MacOS dragging helper☆12Mar 31, 2024Updated 2 years ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- Official code for **Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity** (PruneSI…☆14Mar 25, 2026Updated 6 months ago