Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity simulation, Godzilla KVarN/TriAttention with SM86/SM89 qualification, CUDA weight sharing, and multi-GPU planning.
☆26Aug 22, 2026Updated 3 weeks ago
Alternatives and similar repositories for multi-turboquant
Users that are interested in multi-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆33Mar 2, 2026Updated 6 months ago
- DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers☆372Sep 10, 2026Updated last week
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆54Jul 2, 2026Updated 2 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆19Aug 10, 2026Updated last month
- Open format for specifying structured assumptions and requirements about code.☆74Sep 11, 2026Updated last week
- Tensara's GPU programming problems☆21Apr 23, 2026Updated 4 months ago
- LLM inference in C/C++☆35Aug 24, 2026Updated 3 weeks ago
- Public repo for kernel writing skills in CuTeDSL, Triton, Tilelang and CUDA☆25Jul 7, 2026Updated 2 months ago
- Run `publint` to lint npm packages after the build.☆22Updated this week
- Simple GUI editor for Ideogram V4's structured JSON prompts☆28Jun 10, 2026Updated 3 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- a significantly better pdf and EPUB reader experience for developers☆17Apr 10, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- RWKV-7 mini☆12Mar 29, 2025Updated last year
- SGLang is a fast serving framework for large language models and vision language models.☆34Updated this week
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalit…☆33Jul 23, 2026Updated last month
- Intellisense for package managers files☆17May 31, 2026Updated 3 months ago
- This project uses artificial intelligence technology to analyze video. Recognize video and audio for fragmentation into multiple clip sce…☆11Oct 3, 2018Updated 7 years ago
- Auditable js-to-wasm compiler, focusing on ultra-high performance & security☆26Sep 8, 2026Updated last week
- Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval☆15Jun 24, 2025Updated last year
- Fody AddIn that allows declarative padding structures and classes to fight the false sharing problem.☆12May 24, 2016Updated 10 years ago
- This includes 2 separate tutorial series for OpenAI swarm library each 10 files from basic to advanced☆14Jan 14, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- AGI runtime for bounded recursive self-awareness. Attempting machine consciousness at the Gödel–Turing–Hofstadter Nexus.☆17Updated this week
- Elasticsearch provider for Examine in Umbraco v8☆12Jan 15, 2024Updated 2 years ago
- Extract structured data from documents, images, audio, and video using LLMs in Python☆18Updated this week
- MCP server for Semantic Scholar to search for papers☆20Apr 5, 2025Updated last year
- small board to convert linear 3 pin regulator to switching regulator (TO220 compatible, eg 7805)☆11Oct 8, 2015Updated 10 years ago
- A content-based 3D shape retrieval system that, given a 3D shape, finds the most similar shapes in a given 3D shape database build for th…☆11Nov 10, 2019Updated 6 years ago
- This is my speaker recognition implementation based on the x-vector system described in "X-Vectors: Robust DNN Embeddings for Speaker Rec…☆11Jul 23, 2026Updated last month
- Assembler written in C# for 6502 based systems☆10Dec 17, 2022Updated 3 years ago
- CNTK implementation of Fully Convolutional Networks (FCN) with ResNet for semantic segmentation☆12Aug 18, 2017Updated 9 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Summon React components on demand☆28Sep 26, 2025Updated 11 months ago
- ☆15Feb 15, 2017Updated 9 years ago
- ☆397Apr 16, 2026Updated 5 months ago
- Static analysis for Effect-TS code. Analyze Effect code to extract structure, calculate complexity, and generate visualizations.☆26Updated this week
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 2 months ago
- A template for Zig projects☆15Apr 17, 2026Updated 5 months ago
- Semantic Kernel connector for ONNX models.☆12Jun 10, 2024Updated 2 years ago