Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).
☆23Jun 9, 2026Updated 2 months ago
Alternatives and similar repositories for IceCache
Users that are interested in IceCache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- Segmented Code Adjustment Quantization (SAQ)☆27Sep 22, 2025Updated 11 months ago
- 🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs☆32Aug 14, 2026Updated last week
- DMax: Aggressive Parallel Decoding for dLLMs☆127Jul 5, 2026Updated last month
- RL models to play Sokoban. The fastest recipe wins.☆31Jul 25, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"☆16Apr 30, 2025Updated last year
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆25Dec 5, 2025Updated 8 months ago
- ☆17Mar 20, 2026Updated 5 months ago
- GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent [ICML 2026]☆39Updated this week
- [ICML 2026] Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decoding☆25Aug 9, 2026Updated 2 weeks ago
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆15Aug 16, 2026Updated last week
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 2 months ago
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 2 years ago
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15May 19, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 4 months ago
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 4 months ago
- Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)☆25May 29, 2026Updated 2 months ago
- [ICML 2026] An Evaluation Suite for Chain-of-Thought Controllability☆52Mar 10, 2026Updated 5 months ago
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆25Jul 17, 2026Updated last month
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- A code base for Vexless☆17Mar 7, 2024Updated 2 years ago
- ☆28Mar 18, 2026Updated 5 months ago
- [CVPR 2026 Highlight] PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning☆33Apr 9, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- Official Code Repository for OmniRetrieval☆34Jun 1, 2026Updated 2 months ago
- High-performance K-line (Candlestick) chart for React Native, powered by Skia. Smooth, customizable, and built for real trading apps.☆19Mar 20, 2026Updated 5 months ago
- Query-Adaptive Vector Search☆77Mar 19, 2026Updated 5 months ago
- ☆10Apr 26, 2023Updated 3 years ago
- Public repository for CleANN, an efficient fully-dynamic approximate nearest neighbor search index☆16Sep 18, 2025Updated 11 months ago
- Repository for SIEVE: Effective Filtered Vector Search with Collection of Indexes VLDB '25 submission☆17Jul 30, 2025Updated last year
- ☆33Oct 15, 2025Updated 10 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- ☆10Aug 16, 2024Updated 2 years ago
- Information relating to topics on Data Engineering, Data Infrastructure, Data Storing, Data Warehouses and Business Analysis. For those i…☆13Aug 8, 2021Updated 5 years ago
- ☆19Oct 13, 2025Updated 10 months ago
- Simple and Ideal Circuit Simulation☆13Dec 4, 2017Updated 8 years ago
- The core architecture of the Sovereign Engine. The biological runtime for autonomous organisms.☆23Aug 6, 2026Updated 2 weeks ago
- Docker container for use with Multi-layer Recurrent Neural Networks (LSTM, GRU, RNN) for character-level language models in Torch☆10Jan 19, 2019Updated 7 years ago