Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded deep-context decode (w/ long-range recall) on RTX 4090
☆92Aug 26, 2026Updated this week
Alternatives and similar repositories for llama.cpp-turboq-mtp
Users that are interested in llama.cpp-turboq-mtp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆155Jun 13, 2026Updated 2 months ago
- LLM inference in C/C++☆55Updated this week
- llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% t…☆363Aug 6, 2026Updated 3 weeks ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆985Updated this week
- LLM inference in C/C++☆2,332Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM inference in C/C++☆63May 7, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120…☆36Apr 5, 2026Updated 4 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆227May 14, 2026Updated 3 months ago
- LLM inference in C/C++☆35Updated this week
- Ultra-low-latency LLM gateway with microsecond caching, dynamic routing, budgets, analytics, and forecasting.☆17Apr 2, 2026Updated 4 months ago
- NEW ROCmfp4 format for llama.cpp☆151Jun 13, 2026Updated 2 months ago
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆27May 8, 2026Updated 3 months ago
- ☆12Feb 19, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Experimental LLM inference in C/C++☆40May 15, 2026Updated 3 months ago
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆451Updated this week
- Delegating task/subagent extension for Pi.☆56Updated this week
- ComfyUI implementation of Long-CLIP, including Flux node: LongCLIPTextEncodeFlux☆57Oct 22, 2024Updated last year
- Usage indicator extension for pi with footer status bars and /usage command☆14Apr 17, 2026Updated 4 months ago
- C# DDE Client for MetaTrader 4 (via Ndde)☆10Jan 1, 2018Updated 8 years ago
- A fast, lightweight HTML to Markdown converter optimized for LLM consumption. Uses proven parsing libraries to deliver clean, well-struct…☆88Updated this week
- Pure-Rust LLM runtime: one binary serves (OpenAI-compatible), runs local agents, and distills models on their own rollouts — on Apple Sil…☆18Updated this week
- ☆10Jan 29, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Multimodal AI studio powered by Qwen3.6-35B-A3B. End-to-end web app exposing visual reasoning, image captioning, and document understandi…☆28Apr 23, 2026Updated 4 months ago
- Latent Embedding Operating System☆24Apr 30, 2026Updated 4 months ago
- fast-embeddings-api☆16Nov 23, 2023Updated 2 years ago
- Visualization Library for Python☆27Aug 27, 2019Updated 7 years ago
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆104Aug 22, 2026Updated last week
- A markdown web renderer for AI agents — see the web without screenshots☆67Updated this week
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆50Jul 2, 2026Updated last month
- Deploy Apollo HF space locally☆40Dec 16, 2024Updated last year
- AutoMapper website☆14Aug 19, 2020Updated 6 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- 一款开源的代码管理工具 - 服务端☆19Oct 20, 2025Updated 10 months ago
- ☆45Mar 3, 2026Updated 5 months ago
- 一个语音识别项目☆52May 13, 2025Updated last year
- cursor 0.44.11 download url☆12Feb 7, 2025Updated last year
- 关于Multicharts程序化交易的基础代码(画图,交易,打印输出等)☆10May 21, 2019Updated 7 years ago
- LLM inference in C/C++☆469Updated this week
- AI Avatar 一键启动项目 - 集成 HeyGem 数字人、IndexTTS 语音合成和智能文案生成☆41Jun 9, 2026Updated 2 months ago