KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
☆452Jun 22, 2026Updated last month
Alternatives and similar repositories for KVarN
Users that are interested in KVarN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa☆27Jun 1, 2026Updated 2 months ago
- Welcome to the official repository of AC-LORA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs, a mechanism that provides tr…☆21Nov 14, 2025Updated 8 months ago
- ☆18Sep 26, 2025Updated 10 months ago
- Spire is a Python embedded domain-specific language (DSL) for RTL generation. Its built-in optimizations reduce area and delay of circuit…☆28Aug 3, 2026Updated last week
- Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to make any Large Language Model …☆627May 8, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Fast, lossless LLM inference via dual-view diffusion decoding.☆472May 18, 2026Updated 2 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆852Updated this week
- ☆36Jul 13, 2026Updated 3 weeks ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆31Apr 16, 2026Updated 3 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,840Updated this week
- Efficient Decision tree Ensembles library for IoT edge nodes☆16Jan 29, 2025Updated last year
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,579May 10, 2026Updated 3 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Jul 18, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆110Jun 18, 2026Updated last month
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,921Updated this week
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,915Jul 9, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,024Updated this week
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆3,402Updated this week
- Messy is an open-source framework that integrates a RISC-V ISS with SystemC-AMS☆24Sep 8, 2025Updated 11 months ago
- Riscrithm is a lightweight, low-boilerplate macro-assembly dialect that compiles straight down to pure, human-readable RISC-V assembly. I…☆24Jul 20, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Experimental llama.cpp fork for inference research and development☆731Updated this week
- The inference engine the open-source world built for itself.☆153Aug 2, 2026Updated last week
- Repository for paper: Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets☆27Apr 27, 2026Updated 3 months ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆176Updated this week
- A Plug-and-play Lightweight tool for the Inference Optimization of Deep Neural networks☆52May 8, 2026Updated 3 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,093Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆320Apr 19, 2026Updated 3 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆6,992Updated this week
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,044Apr 23, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆15Feb 21, 2026Updated 5 months ago
- A physics-grounded, cost-aware optimization loop for vLLM☆60Updated this week
- Merkle tree implementation in Rust with configurable storage backends and hash functions. Fixed depth and incremental only. Optimized for…☆229Updated this week
- Jacquard is a small programming language designed for a regime in which most code is written by machine-learning models and reviewed by p…☆114Updated this week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆21,068Updated this week
- Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM☆1,038Aug 4, 2026Updated last week
- Pure Rust Inference Engine☆638Updated this week