☆200Apr 5, 2026Updated 3 months ago
Alternatives and similar repositories for turboquant-model
Users that are interested in turboquant-model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,682Mar 27, 2026Updated 3 months ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago
- ♡☆22Jun 25, 2026Updated 3 weeks ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,031Apr 23, 2026Updated 2 months ago
- # TurboQuant v3 (INT4 + AWQ + Protected Channels + Low-Rank) This notebook demonstrates a **TurboQuant-like** quantization algorithm: - G…☆18Mar 29, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆394Apr 26, 2026Updated 2 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,037Apr 23, 2026Updated 2 months ago
- LLM inference in C/C++☆2,152Updated this week
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆18Mar 28, 2026Updated 3 months ago
- The Screenplay Generator is a web application that allows users to generate a TV or movie screenplay based on a scene template.☆17Mar 18, 2023Updated 3 years ago
- Ultra-Sparse Adaptation of 1-Bit LLMs via XOR Patches☆84Apr 10, 2026Updated 3 months ago
- An extension for oobabooga/text-generation-webui that automatically unloads and reloads your model.☆17Apr 22, 2024Updated 2 years ago
- OpenMOSS pure C++ pipeline based on GGML☆63Jul 11, 2026Updated last week
- interactive semantic search demo using Qwen3-0.6B-Embedding in your browser☆60Feb 25, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- vLLM TurboQuant☆610Jun 25, 2026Updated 3 weeks ago
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆29May 31, 2026Updated last month
- LLM inference in C/C++☆56Updated this week
- rvLLM for runpod serverless environment — lightweight, instant startup vLLM replacement☆40Apr 1, 2026Updated 3 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆21Jun 23, 2026Updated 3 weeks ago
- A FastAPI-based TTS server that provides OpenAI-compatible API endpoints using KittenTTS☆17Aug 16, 2025Updated 11 months ago
- Pure PyTorch implementation of HISA (Hierarchical Indexed Sparse Attention) as a drop-in plugin for HuggingFace models☆16Apr 7, 2026Updated 3 months ago
- Automated generation of comprehensive Agents.md for LLMs, driven by the DSPy Recursive language model implementation.☆252Mar 3, 2026Updated 4 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆17Mar 20, 2026Updated 4 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Just code, nothing else. A community VS Code chat extension on the Claude Agent SDK☆16Updated this week
- LLM inference in C/C++☆36Apr 12, 2026Updated 3 months ago
- Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU pl…☆24Jul 11, 2026Updated last week
- Code Mode inspired local sandboxed MCP Gateway - collapses N servers x M tools into 2 tools (~1,000 tokens)☆149May 14, 2026Updated 2 months ago
- ☆180Aug 10, 2025Updated 11 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week
- This repository is a CUA (computer use agent) system that, using the Qwen3-VL model on Ubuntu computers, aims to perform tasks on your be…☆23Mar 3, 2026Updated 4 months ago
- ☆17Feb 12, 2025Updated last year
- ☆6,997Jun 26, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆483Updated this week
- AI-powered Git CLI tool for smart commits, branches, and PRs☆33Nov 29, 2025Updated 7 months ago
- OpenUI Generative UI tool for Open WebUI - renders interactive charts, forms, tables, and cards in chat via OpenUI Lang.☆58May 25, 2026Updated last month
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 3 months ago
- An AI bot which helps in day-to-day work☆18Apr 7, 2026Updated 3 months ago
- The highest-scoring AI memory system ever benchmarked that isn't reliant on LLM reranking. And it's free & burns less tokens.