Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).
β20Jun 9, 2026Updated last month
Alternatives and similar repositories for IceCache
Users that are interested in IceCache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of NOSAβ20Jun 11, 2026Updated last month
- π₯ [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMsβ30Jun 29, 2026Updated last month
- β16Jul 22, 2026Updated last week
- DMax: Aggressive Parallel Decoding for dLLMsβ128Jul 5, 2026Updated 3 weeks ago
- RL models to play Sokoban. The fastest recipe wins.β31Jul 25, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- β13Jul 15, 2024Updated 2 years ago
- β17Mar 20, 2026Updated 4 months ago
- GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent [ICML 2026]β38Updated this week
- [ICML 2026] Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decodingβ24Jul 20, 2026Updated 2 weeks ago
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! βοΈβοΈβοΈ https://huggingface.co/goodman2001/β¦β15Jul 4, 2026Updated 3 weeks ago
- Repository for "Training Language Models To Explain Their Own Computations"β28Jul 7, 2026Updated 3 weeks ago
- Code for "Reversal Q-Learning (RQL)" for Flow RL from Prior Dataβ34Jun 17, 2026Updated last month
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.β19Jun 12, 2026Updated last month
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inferenceβ15May 19, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Generate a llama-quantize command to copy the quantization parameters of any GGUFβ35Apr 20, 2026Updated 3 months ago
- TurboQuant reference implementation β KV cache compression with engineering insights (ICLR 2026 paper reproduction)β17Mar 28, 2026Updated 4 months ago
- Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)β23May 29, 2026Updated 2 months ago
- β15Jan 22, 2026Updated 6 months ago
- Carla worldsim envrionment for decision-making evaluation and RLβ19Apr 22, 2026Updated 3 months ago
- π Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replayβ24Jul 17, 2026Updated 2 weeks ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.β18Mar 20, 2026Updated 4 months ago
- A code base for Vexlessβ17Mar 7, 2024Updated 2 years ago
- β28Mar 18, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2026 Highlight] PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinningβ33Apr 9, 2026Updated 3 months ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAMβ17May 6, 2026Updated 2 months ago
- High-performance K-line (Candlestick) chart for React Native, powered by Skia. Smooth, customizable, and built for real trading apps.β19Mar 20, 2026Updated 4 months ago
- Public repository for CleANN, an efficient fully-dynamic approximate nearest neighbor search indexβ16Sep 18, 2025Updated 10 months ago
- A cross-modal vector index with fast construction on heterogeneous CPU-GPU environment. Published on DaMoN@SIGMOD 2025.β16Jul 16, 2025Updated last year
- Repository for SIEVE: Effective Filtered Vector Search with Collection of Indexes VLDB '25 submissionβ17Jul 30, 2025Updated last year
- β33Oct 15, 2025Updated 9 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β145Apr 15, 2026Updated 3 months ago
- MacOS dragging helperβ12Mar 31, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.β39Mar 5, 2026Updated 4 months ago
- A toolkit for embedding text datasets with sparse autoencodersβ30Mar 24, 2026Updated 4 months ago
- [VLDB 26, NeurIPS 25] Scalable long-context LLM decoding that leverages sparsityβby treating the KV cache as a vector storage system.β149Updated this week
- β10Aug 16, 2024Updated last year
- Information relating to topics on Data Engineering, Data Infrastructure, Data Storing, Data Warehouses and Business Analysis. For those iβ¦β13Aug 8, 2021Updated 4 years ago
- Implementations of UUID v4 and v7 as defined in the lastest RFC4122 draft. Including a highly-performant custom UUID v8 implementation.β12May 27, 2025Updated last year
- Multimodal Open Source Framework for Conversational Agent Research and Development.β27Feb 16, 2025Updated last year