KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
☆440Jun 22, 2026Updated 3 weeks ago
Alternatives and similar repositories for KVarN
Users that are interested in KVarN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the Paper GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units☆23May 22, 2026Updated last month
- Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa☆27Jun 1, 2026Updated last month
- Welcome to the official repository of AC-LORA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs, a mechanism that provides tr…☆21Nov 14, 2025Updated 8 months ago
- Spire is a Python embedded domain-specific language (DSL) for RTL generation. Its built-in optimizations reduce area and delay of circuit…☆24Updated this week
- Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to make any Large Language Model …☆625May 8, 2026Updated 2 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Fast, lossless LLM inference via dual-view diffusion decoding.☆460May 18, 2026Updated 2 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆789Updated this week
- ☆35Jul 13, 2026Updated last week
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆31Apr 16, 2026Updated 3 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,638Updated this week
- Efficient Decision tree Ensembles library for IoT edge nodes☆16Jan 29, 2025Updated last year
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,500May 10, 2026Updated 2 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆29Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆98Jun 18, 2026Updated last month
- 26m agentic model for tiny devices☆3,225Updated this week
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,702Jul 9, 2026Updated last week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,751Updated this week
- Riscrithm is a lightweight, low-boilerplate macro-assembly dialect that compiles straight down to pure, human-readable RISC-V assembly. I…☆24Updated this week
- LLAMA Turboquant implementation with CUDA support☆703Updated this week
- A collection of optimal and heuristic scheduling tools☆17Apr 24, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Home of ALP/GraphBLAS and ALP/Pregel, featuring shared- and distributed-memory auto-parallelisation of linear algebraic and vertex-centri…☆33Apr 2, 2026Updated 3 months ago
- Repository for paper: Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets☆27Apr 27, 2026Updated 2 months ago
- The inference engine the open-source world built for itself.☆153Jun 13, 2026Updated last month
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆163Jun 27, 2026Updated 3 weeks ago
- Experimental LLM inference in C/C++☆39May 15, 2026Updated 2 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆10,732Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 3 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆6,871Updated this week
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,037Apr 23, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆15Feb 21, 2026Updated 5 months ago
- A physics-grounded, cost-aware optimizer for vLLM.☆54Updated this week
- Merkle tree implementation in Rust with configurable storage backends and hash functions. Fixed depth and incremental only. Optimized for…☆228Jun 15, 2026Updated last month
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆18,904Updated this week
- Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM☆938Jul 2, 2026Updated 2 weeks ago
- PostgreSQL in-database durable execution☆2,659Updated this week
- Pure Rust Inference Engine☆606Updated this week