TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)
☆17Mar 28, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant
Users that are interested in turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML'26] LEMUR reduces multi-vector retrieval for late interaction models such as ColBERT into regular single-vector retrieval.☆33Updated this week
- Segmented Code Adjustment Quantization (SAQ)☆27Sep 22, 2025Updated 11 months ago
- ☆13Jul 15, 2024Updated 2 years ago
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆16Aug 16, 2026Updated last month
- 🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs☆36Aug 14, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Generate fixed dimensional embeddings for multi-dimensional vectors in python based on Muvera from Google.☆21Jun 28, 2025Updated last year
- Implementation of the paper: Selective_Backpropagation from paper Accelerating Deep Learning by Focusing on the Biggest Losers☆15Feb 2, 2020Updated 6 years ago
- Model Analyzer is the Network Statistic Information tool☆16Jul 16, 2026Updated 2 months ago
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 3 years ago
- Faster and Lighter LoRA Implementations☆13Nov 21, 2024Updated last year
- The official implementation of NOSA☆19Jun 11, 2026Updated 3 months ago
- Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).☆25Jun 9, 2026Updated 3 months ago
- ☆21Apr 20, 2026Updated 4 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆25Jun 23, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🦛 Chonkie's recipes for Agents, Context Engineering, and more! 🧑🍳 Chonkie knows how to cook (and it cooks well!)☆18Feb 28, 2026Updated 6 months ago
- BitPolar: near-optimal vector quantization — 3-8 bit compression with zero training. 58 integrations across every major AI framework.☆20Jul 14, 2026Updated 2 months ago
- Supercharge your coding agent with official SurrealDB Skills☆25Updated this week
- ⚡ PDX: A Library for Fast Vector Search and Indexing on CPUs (x86, ARM) — for Python and C++. Index millions of vectors in seconds. Searc…☆97Updated this week
- PersoSim - the open source eID simulator☆16May 21, 2026Updated 3 months ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆36Aug 14, 2024Updated 2 years ago
- SelectiveBackprop accelerates training by dynamically prioritizing useful examples with high loss☆32Mar 12, 2020Updated 6 years ago
- Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development☆31May 7, 2026Updated 4 months ago
- [NeurIPS 2025] Official PyTorch implementation of paper "Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression".☆16Oct 24, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression (DAC'25)☆32Feb 26, 2026Updated 6 months ago
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆21Jul 26, 2026Updated last month
- A custom Arch Linux based kernel for the RoPieee☆11Oct 17, 2021Updated 4 years ago
- Query-Adaptive Vector Search☆77Mar 19, 2026Updated 6 months ago
- Secure (ST31) SDK for Ledger Blue☆14Jul 27, 2018Updated 8 years ago
- NodeJS bitcoin.de API☆10Jun 25, 2018Updated 8 years ago
- RPG^2 is a pure-software system that operates on running C/C++ programs, profiling them, injecting prefetch instructions, and then tuning…☆14May 15, 2024Updated 2 years ago
- 🏆 The winner code for Neurips'23 BigANN Competition OOD and Sparse track.☆15Jun 17, 2025Updated last year
- The Farm-SVE package provides a header that implements the ARM C language extensions (ACLE) for the ARM Scalable Vector Extension (SVE) i…☆15Jan 17, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆30Aug 28, 2023Updated 3 years ago
- ☆37Dec 31, 2025Updated 8 months ago
- ☆104Jul 4, 2025Updated last year
- ☆34Mar 12, 2026Updated 6 months ago
- A DICOM Docker stack with ORTHANC Server, MariaDB database and MedDream frontend viewer.☆13Nov 3, 2022Updated 3 years ago
- ☆39Jul 15, 2025Updated last year
- ☆12Nov 15, 2023Updated 2 years ago