KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
☆497Jun 22, 2026Updated 2 months ago
Alternatives and similar repositories for KVarN
Users that are interested in KVarN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the Paper GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units☆24May 22, 2026Updated 3 months ago
- Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa☆27Jun 1, 2026Updated 3 months ago
- Welcome to the official repository of AC-LORA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs, a mechanism that provides tr…☆21Nov 14, 2025Updated 10 months ago
- ☆18Sep 26, 2025Updated 11 months ago
- Spire is a Python embedded domain-specific language (DSL) for RTL generation. Its built-in optimizations reduce area and delay of circuit…☆39Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to make any Large Language Model …☆630May 8, 2026Updated 4 months ago
- Fast, lossless LLM inference via dual-view diffusion decoding.☆481Aug 12, 2026Updated last month
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,101Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,868Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆32Apr 16, 2026Updated 5 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,151Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,105Aug 18, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,249Updated this week
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆121Jun 18, 2026Updated 3 months ago
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆7,138Jul 9, 2026Updated 2 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,277Updated this week
- Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smar…☆11,854Updated this week
- Experimental llama.cpp fork for inference research and development☆868Updated this week
- Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machi…☆274Updated this week
- Riscrithm is a lightweight, low-boilerplate macro-assembly dialect that compiles straight down to pure, human-readable RISC-V assembly. I…☆27Jul 20, 2026Updated 2 months ago
- A collection of optimal and heuristic scheduling tools☆17Apr 24, 2026Updated 4 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Home of ALP/GraphBLAS and ALP/Pregel, featuring shared- and distributed-memory auto-parallelisation of linear algebraic and vertex-centri…☆33Apr 2, 2026Updated 5 months ago
- The inference engine the open-source world built for itself.☆234Aug 2, 2026Updated last month
- Repository for paper: Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets☆31Apr 27, 2026Updated 4 months ago
- Experimental LLM inference in C/C++☆40May 15, 2026Updated 4 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,875Updated this week
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆225Aug 9, 2026Updated last month
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 5 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆7,060Aug 19, 2026Updated last month
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,049Apr 23, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆17Feb 21, 2026Updated 6 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆22,549Updated this week
- Merkle tree implementation in Rust with configurable storage backends and hash functions. Fixed depth and incremental only. Optimized for…☆228Aug 12, 2026Updated last month
- A physics-grounded, cost-aware optimization loop for vLLM☆67Aug 22, 2026Updated 3 weeks ago
- Jacquard is a small programming language designed for a regime in which most code is written by machine-learning models and reviewed by p…☆119Updated this week
- Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM☆1,123Updated this week
- Pure Rust Inference Engine☆700Updated this week