☆201Apr 5, 2026Updated 4 months ago
Alternatives and similar repositories for turboquant-model
Users that are interested in turboquant-model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,755Updated this week
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- ♡☆23Aug 16, 2026Updated 2 weeks ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,042Apr 23, 2026Updated 4 months ago
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆400Apr 26, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,044Apr 23, 2026Updated 4 months ago
- LLM inference in C/C++☆2,343Updated this week
- The Screenplay Generator is a web application that allows users to generate a TV or movie screenplay based on a scene template.☆18Mar 18, 2023Updated 3 years ago
- Code for the manuscript "A self-supervised multi-layer network of Rectified Spectral Units (ReSUs)" submitted to NeurIPS☆25Apr 6, 2026Updated 4 months ago
- Ultra-Sparse Adaptation of 1-Bit LLMs via XOR Patches☆88Aug 6, 2026Updated 3 weeks ago
- OpenMOSS pure C++ pipeline based on GGML☆72Aug 21, 2026Updated last week
- An extension for oobabooga/text-generation-webui that automatically unloads and reloads your model.☆17Apr 22, 2024Updated 2 years ago
- interactive semantic search demo using Qwen3-0.6B-Embedding in your browser☆60Feb 25, 2026Updated 6 months ago
- vLLM TurboQuant☆618Jun 25, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26Aug 20, 2026Updated 2 weeks ago
- LLM inference in C/C++☆55Updated this week
- A local code review tool designed for the coding agent workflow☆195Jun 23, 2026Updated 2 months ago
- OCTAVE protocol - structured AI communication with 3-20x token reduction. MCP server with lenient-to-canonical pipeline and schema valida…☆55Jun 23, 2026Updated 2 months ago
- rvLLM for runpod serverless environment — lightweight, instant startup vLLM replacement☆40Apr 1, 2026Updated 5 months ago
- A FastAPI-based TTS server that provides OpenAI-compatible API endpoints using KittenTTS☆17Aug 16, 2025Updated last year
- Provides the screen image for multimodal models when you send a message.☆13Nov 1, 2025Updated 10 months ago
- Automated generation of comprehensive Agents.md for LLMs, driven by the DSPy Recursive language model implementation.☆255Mar 3, 2026Updated 6 months ago
- ☆26Mar 18, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- LLM inference in C/C++☆35Aug 24, 2026Updated last week
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆25Aug 22, 2026Updated last week
- Production Claude Code skills for n8n from a Verified Creator's 100+ workflows☆31Apr 26, 2026Updated 4 months ago
- Code Mode inspired local sandboxed MCP Gateway - collapses N servers x M tools into 2 tools (~1,000 tokens)☆151May 14, 2026Updated 3 months ago
- ☆179Aug 10, 2025Updated last year
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week
- This repository is a CUA (computer use agent) system that, using the Qwen3-VL model on Ubuntu computers, aims to perform tasks on your be…☆22Mar 3, 2026Updated 6 months ago
- ☆7,019Jul 20, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆17Feb 12, 2025Updated last year
- DoubleAI’s hyperoptimised version of cuGraph☆65Mar 3, 2026Updated 6 months ago
- ☆17May 16, 2025Updated last year
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆516Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 4 months ago
- VZBot style frame brace for Voron Trident☆12Dec 8, 2022Updated 3 years ago
- ☆16Feb 3, 2026Updated 7 months ago