vLLM TurboQuant
☆619Jun 25, 2026Updated 2 months ago
Alternatives and similar repositories for vllm-turboquant
Users that are interested in vllm-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,777Sep 3, 2026Updated 2 weeks ago
- AI Coding Factory. Replace your software outsourcing vendor contracts with your own AI.☆176Feb 2, 2026Updated 7 months ago
- Aegis- a local zero-trust AI gate for OS and Apps packages☆31May 22, 2026Updated 3 months ago
- Welcome to Firecracker + AgentFS for AI Agents☆31May 18, 2026Updated 4 months ago
- Claude Code native build of `autoresearch`: an autonomous experiment loop that writes session files, benchmarks changes, logs results, an…☆27Mar 13, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,048Apr 23, 2026Updated 4 months ago
- dede (DogEatDog.DependencyExplorer) is a static dependency and blast-radius explorer for mixed .NET workspaces with many sibling reposi…☆33Mar 10, 2026Updated 6 months ago
- Experimental llama.cpp fork for inference research and development☆860Updated this week
- ☆7,030Jul 20, 2026Updated last month
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆471Aug 17, 2026Updated last month
- ☆23Apr 10, 2026Updated 5 months ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,046Apr 23, 2026Updated 4 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- RLM based security scanner for massive .NET codebases☆78Feb 9, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆403Apr 26, 2026Updated 4 months ago
- LLM inference in C/C++☆2,386Sep 7, 2026Updated last week
- ☆592Sep 5, 2026Updated 2 weeks ago
- FTE+AI: Vendor Replacement Program Framework☆82Dec 31, 2025Updated 8 months ago
- SkillOpt with local AI is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-dr…☆85May 25, 2026Updated 3 months ago
- An agent-first dependency and blast-radius explorer for Python codebases. Generates structured, machine-readable dependency graphs that A…☆19Mar 18, 2026Updated 6 months ago
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated 2 months ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆27Jun 16, 2026Updated 3 months ago
- llama-benchy - llama-bench style benchmarking tool for all backends☆718Jul 10, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 2 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,100Aug 18, 2026Updated last month
- SimpleAgents lets anyone vibe-code LLM agents and ship them production-ready. It’s Rust-first with Python/Node/Go bindings, multi-provide…☆33Jun 6, 2026Updated 3 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆104Updated this week
- ☆397Apr 16, 2026Updated 5 months ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆17,197Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,867Updated this week
- ♡☆23Aug 16, 2026Updated last month
- Fabric for Agents☆45Jul 20, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Docker configuration for running VLLM on dual DGX Sparks☆2,285Updated this week
- ☆199Apr 5, 2026Updated 5 months ago
- Recursive Language Models Gateway build using rlm https://arxiv.org/abs/2512.24601v1☆127Jan 6, 2026Updated 8 months ago
- Crack - Make your lid loud!☆16Mar 22, 2026Updated 5 months ago
- ☆156Mar 31, 2026Updated 5 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,560Mar 19, 2026Updated 6 months ago
- An open-source template that turns your repo into an autonomous development team. Uses GitHub Actions and Claude to orchestrate AI agents…☆28Mar 19, 2026Updated 6 months ago