LLM inference in C/C++
☆64May 7, 2026Updated 2 months ago
Alternatives and similar repositories for llama-cpp-turboquant
Users that are interested in llama-cpp-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 3 months ago
- ☆41Mar 31, 2026Updated 4 months ago
- ça Streak et ça Master☆12Updated this week
- LLM inference in C/C++☆2,226Updated this week
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆22Jun 23, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120…☆35Apr 5, 2026Updated 3 months ago
- Unified KV-cache compression toolkit for LLM inference: 12 Python-native methods, Godzilla KVarN/TriAttention + pinned Gigatoken runtime …☆25Updated this week
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆91Updated this week
- Distribute and run LLMs with a single file.☆25May 13, 2025Updated last year
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- Experimental llama.cpp fork for inference research and development☆725Updated this week
- A way to visualize nginx config files☆10Nov 21, 2022Updated 3 years ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆829Updated this week
- llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% t…☆323Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆16Jul 1, 2026Updated last month
- 30-C#-Project☆13Jan 8, 2024Updated 2 years ago
- Simplifies development, reduces developer frustration, and speeds up website creation process.☆14Jan 20, 2023Updated 3 years ago
- LLM inference in C/C++☆442Updated this week
- Car Monitoring System | سیستم نظارت بر خودرو☆15Nov 8, 2024Updated last year
- MCP Think tool prebuilt binaries and code☆18Mar 27, 2025Updated last year
- ☆24Dec 10, 2025Updated 7 months ago
- Founder of omnipkg — Run multiple conflicting Python package & interpreter versions concurrently in one environment in milliseconds. On @…☆18Apr 27, 2026Updated 3 months ago
- ☆34Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- "O(log N) MoE Expert routing via RT Core ray tracing. BVH traversal replaces matrix multiplication in neural language models."☆83Apr 13, 2026Updated 3 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆422Jul 25, 2026Updated last week
- A collection of useful .gitignore templates☆18May 20, 2026Updated 2 months ago
- A framework for efficient model inference with omni-modality models☆29Jul 12, 2026Updated 3 weeks ago
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated 3 weeks ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- Advanced intelligent-automation capabilities converge in NodeSync, optimizing node interactions across a scalable, modern, cloud-agnostic…☆14May 12, 2026Updated 2 months ago
- I love building modern web apps, exploring AI-powered development, and solving real-world problems with code. Currently focusing on MERN …☆22Feb 11, 2026Updated 5 months ago
- Cheat sheets for various commands and scripts☆17Jun 29, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- llama.cpp fork with additional SOTA quants and improved performance☆2,990Updated this week
- A C# .netstandard client library for the Kraken REST and Websocket Spot and Futures API focusing on clear usage and models☆20Apr 7, 2024Updated 2 years ago
- Python Advanced☆18Nov 30, 2023Updated 2 years ago
- Logos is a readable scripting language with C-like syntax, sane error handling, built-in concurrency, and binary compilation.☆15Mar 22, 2026Updated 4 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 4 months ago
- AI-powered domain name generator with real-time WHOIS availability checking☆15Apr 13, 2026Updated 3 months ago
- Parser for Rust source code☆20Dec 19, 2025Updated 7 months ago