TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)
☆17Mar 28, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant
Users that are interested in turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Aug 28, 2025Updated last year
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15May 19, 2026Updated 3 months ago
- [ICML'26] LEMUR reduces multi-vector retrieval for late interaction models such as ColBERT into regular single-vector retrieval.☆33Aug 23, 2026Updated last week
- ☆13Jul 15, 2024Updated 2 years ago
- 🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs☆34Aug 14, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆15Aug 16, 2026Updated last week
- Generate fixed dimensional embeddings for multi-dimensional vectors in python based on Muvera from Google.☆21Jun 28, 2025Updated last year
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 2 years ago
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- A project which does the ColBERT pruning based on the LP or L1 norm☆20Jun 11, 2025Updated last year
- basic api for streaming data from the keyence LJV-7300 Line Scanner☆15Sep 20, 2016Updated 9 years ago
- Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).☆23Jun 9, 2026Updated 2 months ago
- ☆29Jul 25, 2025Updated last year
- ☆19Apr 20, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Modular task agnostic training pipeline using LFM2 from Liquid AI with unsloth.☆16Sep 13, 2025Updated 11 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆24Jun 23, 2026Updated 2 months ago
- 🦛 Chonkie's recipes for Agents, Context Engineering, and more! 🧑🍳 Chonkie knows how to cook (and it cooks well!)☆18Feb 28, 2026Updated 6 months ago
- ☆10Mar 2, 2022Updated 4 years ago
- Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers☆24May 26, 2026Updated 3 months ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated 2 years ago
- miniDSP BEQ Files☆10Nov 7, 2023Updated 2 years ago
- 📑 The best way to have your Agents read the documentation! 🤖☆35Mar 25, 2026Updated 5 months ago
- Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development☆31May 7, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression (DAC'25)☆32Feb 26, 2026Updated 6 months ago
- ☆32Oct 22, 2023Updated 2 years ago
- ☆26Jan 11, 2022Updated 4 years ago
- RPG^2 is a pure-software system that operates on running C/C++ programs, profiling them, injecting prefetch instructions, and then tuning…☆14May 15, 2024Updated 2 years ago
- 🏆 The winner code for Neurips'23 BigANN Competition OOD and Sparse track.☆15Jun 17, 2025Updated last year
- The Farm-SVE package provides a header that implements the ARM C language extensions (ACLE) for the ARM Scalable Vector Extension (SVE) i…☆15Jan 17, 2024Updated 2 years ago
- ☆37Dec 31, 2025Updated 7 months ago
- The Docker image for hickory-dns☆15Apr 16, 2026Updated 4 months ago
- Framework to build general purpose distributed data platform in Azure Virtual Machines to support various platforms like Hadoop, Cassandr…☆20Nov 28, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆104Jul 4, 2025Updated last year
- ☆34Mar 12, 2026Updated 5 months ago
- [SIGMOD 2025] PQCache: Product Quantization-based KVCache for Long Context LLM Inference☆91Dec 7, 2025Updated 8 months ago
- Roaring Bitmap positional phrase matching for low-latency LLM context retrieval.☆29May 4, 2026Updated 3 months ago
- A DICOM Docker stack with ORTHANC Server, MariaDB database and MedDream frontend viewer.☆13Nov 3, 2022Updated 3 years ago
- A list of multi-vector retrieval resources☆18May 29, 2024Updated 2 years ago
- Official repository of the Seismic library.☆136Aug 6, 2026Updated 3 weeks ago