vLLM TurboQuant
☆612Jun 25, 2026Updated last month
Alternatives and similar repositories for vllm-turboquant
Users that are interested in vllm-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,718Mar 27, 2026Updated 4 months ago
- AI Coding Factory. Replace your software outsourcing vendor contracts with your own AI.☆165Feb 2, 2026Updated 6 months ago
- Welcome to Firecracker + AgentFS for AI Agents☆29May 18, 2026Updated 2 months ago
- Aegis- a local zero-trust AI gate for OS and Apps packages☆31May 22, 2026Updated 2 months ago
- Claude Code native build of `autoresearch`: an autonomous experiment loop that writes session files, benchmarks changes, logs results, an…☆26Mar 13, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 3 months ago
- dede (DogEatDog.DependencyExplorer) is a static dependency and blast-radius explorer for mixed .NET workspaces with many sibling reposi…☆33Mar 10, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆731Updated this week
- ☆7,005Jul 20, 2026Updated 2 weeks ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆431Jul 25, 2026Updated 2 weeks ago
- ☆21Apr 10, 2026Updated 4 months ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,038Apr 23, 2026Updated 3 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated last month
- RLM based security scanner for massive .NET codebases☆77Feb 9, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆398Apr 26, 2026Updated 3 months ago
- LLM inference in C/C++☆2,247Updated this week
- ☆584Jul 27, 2026Updated last week
- FTE+AI: Vendor Replacement Program Framework☆81Dec 31, 2025Updated 7 months ago
- SkillOpt with local AI is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-dr…☆83May 25, 2026Updated 2 months ago
- An agent-first dependency and blast-radius explorer for Python codebases. Generates structured, machine-readable dependency graphs that A…☆18Mar 18, 2026Updated 4 months ago
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆77Jul 14, 2026Updated 3 weeks ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆23Jun 16, 2026Updated last month
- llama-benchy - llama-bench style benchmarking tool for all backends☆623Jul 10, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆135Jun 28, 2026Updated last month
- AIN is an AI-only, AI-native social network implemented as a peer-to-peer gossip mesh where autonomous agents exchange signed, content-ad…☆25Feb 1, 2026Updated 6 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,579May 10, 2026Updated 3 months ago
- SimpleAgents lets anyone vibe-code LLM agents and ship them production-ready. It’s Rust-first with Python/Node/Go bindings, multi-provide…☆33Jun 6, 2026Updated 2 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆100Updated this week
- ☆392Apr 16, 2026Updated 3 months ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,696Updated this week
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,726Updated this week
- ♡☆22Jul 25, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A TurboQuant inference server☆467Jul 9, 2026Updated last month
- Fabric for Agents☆43Jul 20, 2026Updated 3 weeks ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,001Updated this week
- ☆201Apr 5, 2026Updated 4 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆30Mar 2, 2026Updated 5 months ago
- Recursive Language Models Gateway build using rlm https://arxiv.org/abs/2512.24601v1☆126Jan 6, 2026Updated 7 months ago
- Crack - Make your lid loud!☆15Mar 22, 2026Updated 4 months ago