☆200Apr 5, 2026Updated 4 months ago
Alternatives and similar repositories for turboquant-model
Users that are interested in turboquant-model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,721Mar 27, 2026Updated 4 months ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- Install, update and remove AppImage from your CLI. appimage, linux, package-manager☆19May 21, 2025Updated last year
- ♡☆22Jul 25, 2026Updated 2 weeks ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,038Apr 23, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆397Apr 26, 2026Updated 3 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,044Apr 23, 2026Updated 3 months ago
- LLM inference in C/C++☆2,250Updated this week
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 4 months ago
- The Screenplay Generator is a web application that allows users to generate a TV or movie screenplay based on a scene template.☆18Mar 18, 2023Updated 3 years ago
- OpenMOSS pure C++ pipeline based on GGML☆67Jul 31, 2026Updated last week
- interactive semantic search demo using Qwen3-0.6B-Embedding in your browser☆60Feb 25, 2026Updated 5 months ago
- vLLM TurboQuant☆612Jun 25, 2026Updated last month
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26May 31, 2026Updated 2 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- LLM inference in C/C++☆56Updated this week
- OCTAVE protocol - structured AI communication with 3-20x token reduction. MCP server with lenient-to-canonical pipeline and schema valida…☆55Jun 23, 2026Updated last month
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆23Jun 23, 2026Updated last month
- rvLLM for runpod serverless environment — lightweight, instant startup vLLM replacement☆40Apr 1, 2026Updated 4 months ago
- Provides the screen image for multimodal models when you send a message.☆13Nov 1, 2025Updated 9 months ago
- Automated generation of comprehensive Agents.md for LLMs, driven by the DSPy Recursive language model implementation.☆253Mar 3, 2026Updated 5 months ago
- ☆26Mar 18, 2026Updated 4 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 4 months ago
- Just code, nothing else. A community VS Code chat extension on the Claude Agent SDK☆16Jul 15, 2026Updated 3 weeks ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, ex…☆25Aug 3, 2026Updated last week
- ☆180Aug 10, 2025Updated last year
- Python SDK for mcpd☆24Feb 24, 2026Updated 5 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week
- ☆7,003Jul 20, 2026Updated 3 weeks ago
- ☆17May 16, 2025Updated last year
- AutoGPT maintainer/reviewer system☆16May 26, 2023Updated 3 years ago
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆509Updated this week
- A skill that convenes a panel of CLI-based AI agents (Claude Code, Codex, Gemini CLI) to deliberate on engineering problems through struc…☆88Apr 7, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- OpenUI Generative UI tool for Open WebUI - renders interactive charts, forms, tables, and cards in chat via OpenUI Lang.☆64May 25, 2026Updated 2 months ago
- ☆28Jul 10, 2026Updated last month
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆320Apr 19, 2026Updated 3 months ago
- ☆16Feb 3, 2026Updated 6 months ago
- The operating layer for AI coding agents.☆17Jun 23, 2026Updated last month
- Voice models for Piper TTS trained using TextyMcSpeechy☆24Jun 29, 2025Updated last year
- Workflow orchestration for AI coding agent swarms, from task to merged PR.☆1,027Updated this week