Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).
☆25Jun 9, 2026Updated 3 months ago
Alternatives and similar repositories for IceCache
Users that are interested in IceCache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of NOSA☆20Jun 11, 2026Updated 3 months ago
- Segmented Code Adjustment Quantization (SAQ)☆27Sep 22, 2025Updated 11 months ago
- 🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs☆35Aug 14, 2026Updated 3 weeks ago
- ☆17Aug 27, 2026Updated 2 weeks ago
- DMax: Aggressive Parallel Decoding for dLLMs☆131Jul 5, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- RL models to play Sokoban. The fastest recipe wins.☆31Jul 25, 2026Updated last month
- ☆13Jul 15, 2024Updated 2 years ago
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆26Dec 5, 2025Updated 9 months ago
- GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent [ICML 2026]☆41Updated this week
- [ICML 2026] Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decoding☆27Aug 9, 2026Updated last month
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆16Aug 16, 2026Updated 3 weeks ago
- Repository for "Training Language Models To Explain Their Own Computations"☆37Jul 7, 2026Updated 2 months ago
- ☆12May 19, 2025Updated last year
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Jan 7, 2025Updated last year
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15May 19, 2026Updated 3 months ago
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 4 months ago
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 5 months ago
- Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)☆26May 29, 2026Updated 3 months ago
- ☆14Jan 22, 2026Updated 7 months ago
- Carla worldsim envrionment for decision-making evaluation and RL☆19Apr 22, 2026Updated 4 months ago
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆25Jul 17, 2026Updated last month
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official Code Repository for OmniRetrieval☆35Jun 1, 2026Updated 3 months ago
- ☆17Mar 24, 2025Updated last year
- ☆10Apr 26, 2023Updated 3 years ago
- Repository for SIEVE: Effective Filtered Vector Search with Collection of Indexes VLDB '25 submission☆17Jul 30, 2025Updated last year
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 4 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- Official code for **Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity** (PruneSI…☆13Mar 25, 2026Updated 5 months ago
- A toolkit for embedding text datasets with sparse autoencoders☆31Mar 24, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model☆16May 27, 2025Updated last year
- [VLDB 26, NeurIPS 25] Scalable long-context LLM decoding that leverages sparsity—by treating the KV cache as a vector storage system.☆152Jul 30, 2026Updated last month
- ☆10Aug 16, 2024Updated 2 years ago
- Information relating to topics on Data Engineering, Data Infrastructure, Data Storing, Data Warehouses and Business Analysis. For those i…☆13Aug 8, 2021Updated 5 years ago
- ☆19Oct 13, 2025Updated 11 months ago
- Implementations of UUID v4 and v7 as defined in the lastest RFC4122 draft. Including a highly-performant custom UUID v8 implementation.☆12May 27, 2025Updated last year
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆178Feb 27, 2026Updated 6 months ago