☆199Apr 5, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant-model
Users that are interested in turboquant-model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,784Sep 3, 2026Updated 2 weeks ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,046Apr 23, 2026Updated 5 months ago
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆402Apr 26, 2026Updated 4 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,049Apr 23, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LLM inference in C/C++☆2,397Updated this week
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 5 months ago
- The Screenplay Generator is a web application that allows users to generate a TV or movie screenplay based on a scene template.☆19Mar 18, 2023Updated 3 years ago
- Code for the manuscript "A self-supervised multi-layer network of Rectified Spectral Units (ReSUs)" submitted to NeurIPS☆25Apr 6, 2026Updated 5 months ago
- Ultra-Sparse Adaptation of 1-Bit LLMs via XOR Patches☆91Aug 6, 2026Updated last month
- OpenMOSS pure C++ pipeline based on GGML☆72Aug 21, 2026Updated last month
- An extension for oobabooga/text-generation-webui that automatically unloads and reloads your model.☆17Apr 22, 2024Updated 2 years ago
- interactive semantic search demo using Qwen3-0.6B-Embedding in your browser☆60Feb 25, 2026Updated 6 months ago
- vLLM TurboQuant☆618Jun 25, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26Aug 20, 2026Updated last month
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆25Jun 23, 2026Updated 3 months ago
- rvLLM for runpod serverless environment — lightweight, instant startup vLLM replacement☆40Apr 1, 2026Updated 5 months ago
- python script able to connect to the chessnut air board using bluetooth☆18Dec 30, 2022Updated 3 years ago
- LLM inference in C/C++☆35Aug 24, 2026Updated 3 weeks ago
- Just code, nothing else. A community VS Code chat extension on the Claude Agent SDK☆17Updated this week
- Production Claude Code skills for n8n from a Verified Creator's 100+ workflows☆32Apr 26, 2026Updated 4 months ago
- ☆17Dec 29, 2025Updated 8 months ago
- Code Mode inspired local sandboxed MCP Gateway - collapses N servers x M tools into 2 tools (~1,000 tokens)☆151Sep 10, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆179Aug 10, 2025Updated last year
- This repository is a CUA (computer use agent) system that, using the Qwen3-VL model on Ubuntu computers, aims to perform tasks on your be…☆22Mar 3, 2026Updated 6 months ago
- ☆17Feb 12, 2025Updated last year
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆168Sep 11, 2026Updated last week
- DoubleAI’s hyperoptimised version of cuGraph☆67Mar 3, 2026Updated 6 months ago
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆521Sep 15, 2026Updated last week
- A skill that convenes a panel of CLI-based AI agents (Claude Code, Codex, Gemini CLI) to deliberate on engineering problems through struc…☆90Apr 7, 2026Updated 5 months ago
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 5 months ago
- VZBot style frame brace for Voron Trident☆12Dec 8, 2022Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An AI bot which helps in day-to-day work☆18Apr 7, 2026Updated 5 months ago
- AI-powered work journal template for software developers. PARA structure, Bullet Journal daily rhythm, Zettelkasten knowledge graph with …☆40May 21, 2026Updated 4 months ago
- The highest-scoring AI memory system ever benchmarked that isn't reliant on LLM reranking. And it's free & burns less tokens.☆36Jun 1, 2026Updated 3 months ago
- Just an UI for Chatterbox, which uses about 1-2 GB RAM. Double click and you're good to go.☆24Jun 11, 2026Updated 3 months ago
- ☆27Sep 25, 2024Updated last year
- Tensor parallelism is all you need. Run LLMs on an AI cluster at home using any device. Distribute the workload, divide RAM usage, and in…☆18Nov 11, 2024Updated last year
- Python virtual environment manager for @xonsh.☆31Aug 31, 2026Updated 3 weeks ago