TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
☆77Jul 14, 2026Updated 3 weeks ago
Alternatives and similar repositories for turboquant-vllm
Users that are interested in turboquant-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- ☆27Mar 31, 2026Updated 4 months ago
- ☆21Mar 28, 2026Updated 4 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- NVIDIA Linux open GPU with P2P support☆377Jul 13, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A simple GPT-3 interface to automate core legal writing tasks☆13Mar 8, 2023Updated 3 years ago
- HTTP Proxy that allows you to define the IP address at connect time☆20Jan 11, 2026Updated 6 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆23May 5, 2026Updated 2 months ago
- OpenCode port of Flow-Next: plan-first workflows, Ralph autonomous mode (overnight coding with fresh context), multi-model review gates v…☆39Jan 23, 2026Updated 6 months ago
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated last month
- StarCraft BroodWar Hacker Finder, anti-hack, replay analyzer-organizer and utility tool☆10Nov 1, 2023Updated 2 years ago
- A Hermes UI☆23Jun 6, 2026Updated last month
- Electron speech-to-speech app for your voice calls based on 100% locally run AI models☆35Jul 23, 2025Updated last year
- ☆60Feb 24, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Developing a legal research tool leveraging ChatGPT / GPT-4☆14Mar 10, 2024Updated 2 years ago
- vLLM TurboQuant☆611Jun 25, 2026Updated last month
- Low-latency IPC library for building persistent agentic tool servers (LLM inference, TTS, vector search, browser automation) over named p…☆15Jun 27, 2026Updated last month
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆34Jan 6, 2026Updated 6 months ago
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse☆131Mar 14, 2026Updated 4 months ago
- Simple intermediate representation language for learning and research.☆22Mar 27, 2020Updated 6 years ago
- Experimental llama.cpp fork for inference research and development☆725Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆2,990Updated this week
- Graph QABot Demo| 图谱问答案例☆14Apr 11, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- LLM inference in C/C++☆35Apr 12, 2026Updated 3 months ago
- An interactive course on prompt engineering for lawyers built using streamlit.☆19Sep 1, 2024Updated last year
- Replayable Browser Agent☆16Apr 24, 2026Updated 3 months ago
- "rsync for cloud storage" - Google Drive, S3, Dropbox, Backblaze B2, One Drive, Swift, Hubic, Wasabi, Google Cloud Storage, Yandex Files☆11Oct 27, 2024Updated last year
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆98Updated this week
- ☆20Oct 25, 2025Updated 9 months ago
- ☆10Feb 9, 2025Updated last year
- A Triton JIT runtime and ffi provider in C++☆38Jul 28, 2026Updated last week
- An Opencode plugin for managing git worktrees.☆77Mar 25, 2026Updated 4 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Enhancing the convergence speed by 2x and improving the training success of Physics-Informed Neural Networks (PINNs).☆13Oct 14, 2024Updated last year
- Run DeepSeek-V4-Flash on SM89 (Ada / RTX 4090) with vLLM — patch over PR #41834. Validated on 4x RTX 4090.☆51Updated this week
- 中国法研杯 CAIL 2019☆13Jun 17, 2019Updated 7 years ago
- bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vi…☆15Mar 12, 2026Updated 4 months ago
- a fast and lightweight distributed background task processing framework with seamless scheduling.☆15Mar 30, 2026Updated 4 months ago
- Distributed Inference for mlx LLm☆102Aug 1, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week