KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
☆480Jun 22, 2026Updated 2 months ago
Alternatives and similar repositories for KVarN
Users that are interested in KVarN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the Paper GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units☆23May 22, 2026Updated 3 months ago
- Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa☆27Jun 1, 2026Updated 2 months ago
- Welcome to the official repository of AC-LORA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs, a mechanism that provides tr…☆21Nov 14, 2025Updated 9 months ago
- ☆18Sep 26, 2025Updated 11 months ago
- Spire is a Python embedded domain-specific language (DSL) for RTL generation. Its built-in optimizations reduce area and delay of circuit…☆38Aug 17, 2026Updated 2 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to make any Large Language Model …☆629May 8, 2026Updated 3 months ago
- Fast, lossless LLM inference via dual-view diffusion decoding.☆477Aug 12, 2026Updated 2 weeks ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆985Updated this week
- ☆36Jul 13, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆31Apr 16, 2026Updated 4 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,042Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,011Aug 18, 2026Updated last week
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Jul 18, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆115Jun 18, 2026Updated 2 months ago
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,122Updated this week
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆7,056Jul 9, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆9,776Updated this week
- Messy is an open-source framework that integrates a RISC-V ISS with SystemC-AMS☆25Sep 8, 2025Updated 11 months ago
- Riscrithm is a lightweight, low-boilerplate macro-assembly dialect that compiles straight down to pure, human-readable RISC-V assembly. I…☆27Jul 20, 2026Updated last month
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A collection of optimal and heuristic scheduling tools☆17Apr 24, 2026Updated 4 months ago
- Experimental llama.cpp fork for inference research and development☆789Updated this week
- Home of ALP/GraphBLAS and ALP/Pregel, featuring shared- and distributed-memory auto-parallelisation of linear algebraic and vertex-centri…☆33Apr 2, 2026Updated 4 months ago
- The inference engine the open-source world built for itself.☆233Aug 2, 2026Updated 3 weeks ago
- Repository for paper: Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets☆27Apr 27, 2026Updated 4 months ago
- Experimental LLM inference in C/C++☆40May 15, 2026Updated 3 months ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆178Aug 9, 2026Updated 3 weeks ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,572Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆7,021Aug 19, 2026Updated last week
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,042Apr 23, 2026Updated 4 months ago
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆16Feb 21, 2026Updated 6 months ago
- A physics-grounded, cost-aware optimization loop for vLLM☆64Aug 22, 2026Updated last week
- Merkle tree implementation in Rust with configurable storage backends and hash functions. Fixed depth and incremental only. Optimized for…☆229Aug 12, 2026Updated 2 weeks ago
- Jacquard is a small programming language designed for a regime in which most code is written by machine-learning models and reviewed by p…☆116Updated this week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆21,927Updated this week