vLLM TurboQuant
☆610Jun 25, 2026Updated 3 weeks ago
Alternatives and similar repositories for vllm-turboquant
Users that are interested in vllm-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,682Mar 27, 2026Updated 3 months ago
- AI Coding Factory. Replace your software outsourcing vendor contracts with your own AI.☆161Feb 2, 2026Updated 5 months ago
- Aegis- a local zero-trust AI gate for OS and Apps packages☆31May 22, 2026Updated last month
- Welcome to Firecracker + AgentFS for AI Agents☆29May 18, 2026Updated 2 months ago
- Claude Code native build of `autoresearch`: an autonomous experiment loop that writes session files, benchmarks changes, logs results, an…☆26Mar 13, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆24Mar 26, 2026Updated 3 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,037Apr 23, 2026Updated 2 months ago
- dede (DogEatDog.DependencyExplorer) is a static dependency and blast-radius explorer for mixed .NET workspaces with many sibling reposi…☆33Mar 10, 2026Updated 4 months ago
- LLAMA Turboquant implementation with CUDA support☆703Updated this week
- ☆6,997Jun 26, 2026Updated 3 weeks ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆396May 29, 2026Updated last month
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,031Apr 23, 2026Updated 2 months ago
- ☆21Apr 10, 2026Updated 3 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆52Jun 28, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- LLM inference in C/C++☆2,152Updated this week
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆394Apr 26, 2026Updated 2 months ago
- RLM based security scanner for massive .NET codebases☆76Feb 9, 2026Updated 5 months ago
- ☆579Jun 26, 2026Updated 3 weeks ago
- FTE+AI: Vendor Replacement Program Framework☆81Dec 31, 2025Updated 6 months ago
- SkillOpt with local AI is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-dr…☆80May 25, 2026Updated last month
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆75Jul 14, 2026Updated last week
- An agent-first dependency and blast-radius explorer for Python codebases. Generates structured, machine-readable dependency graphs that A…☆18Mar 18, 2026Updated 4 months ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆20Jun 16, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- llama-benchy - llama-bench style benchmarking tool for all backends☆580Jul 10, 2026Updated last week
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆127Jun 28, 2026Updated 3 weeks ago
- AIN is an AI-only, AI-native social network implemented as a peer-to-peer gossip mesh where autonomous agents exchange signed, content-ad…☆25Feb 1, 2026Updated 5 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,500May 10, 2026Updated 2 months ago
- SimpleAgents lets anyone vibe-code LLM agents and ship them production-ready. It’s Rust-first with Python/Node/Go bindings, multi-provide…☆34Jun 6, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆96Jul 13, 2026Updated last week
- A vector index built on TurboQuant, written in Rust with Python bindings☆13,639Updated this week
- ☆389Apr 16, 2026Updated 3 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ♡☆22Jun 25, 2026Updated 3 weeks ago
- Docker configuration for running VLLM on dual DGX Sparks☆1,859Updated this week
- ☆200Apr 5, 2026Updated 3 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆25Mar 2, 2026Updated 4 months ago
- Recursive Language Models Gateway build using rlm https://arxiv.org/abs/2512.24601v1☆126Jan 6, 2026Updated 6 months ago
- Crack - Make your lid loud!☆15Mar 22, 2026Updated 3 months ago
- ☆155Mar 31, 2026Updated 3 months ago