Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU.
☆24Jul 11, 2026Updated last week
Alternatives and similar repositories for multi-turboquant
Users that are interested in multi-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago
- LLM inference in C/C++☆64May 7, 2026Updated 2 months ago
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆19Jul 6, 2026Updated 2 weeks ago
- LLM inference in C/C++☆36Apr 12, 2026Updated 3 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆17Mar 20, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Simple GUI editor for Ideogram V4's structured JSON prompts☆25Jun 10, 2026Updated last month
- TDS – Traffic Delivery System☆12Oct 9, 2021Updated 4 years ago
- High-performance K-line (Candlestick) chart for React Native, powered by Skia. Smooth, customizable, and built for real trading apps.☆19Mar 20, 2026Updated 4 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆52Jun 28, 2026Updated 3 weeks ago
- RWKV-7 mini☆12Mar 29, 2025Updated last year
- A2A MCP Server is a lightweight Python bridge that lets Claude Desktop or any MCP client talk to A2A agents. It provides three tools: reg…☆21May 4, 2025Updated last year
- JavaScript component to parse, clean, remove formatting (unformat) numbers in strings.☆10Dec 5, 2024Updated last year
- Dispatch AI coding agents (Claude Code, Copilot, OpenCode) as MCP tools☆21Apr 8, 2026Updated 3 months ago
- ☆12Dec 8, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Manages tasks and persistent storage for AI agents — assigns work, tracks completion, and provides a shared key-value store with full-tex…☆32Updated this week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 3 months ago
- TurboDB — a blazing-fast NoSQL document database written in Zig. mmap + WAL + B-tree + MVCC.☆19May 1, 2026Updated 2 months ago
- A simple NER implementation using a DistilBERT based model with ML.NET☆13May 6, 2021Updated 5 years ago
- Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval