TurboQuant KV cache compression for llama.cpp — HIP/ROCm port for AMD RDNA3 (gfx1100)
☆57Apr 26, 2026Updated 3 months ago
Alternatives and similar repositories for llama.cpp-turboquant-hip
Users that are interested in llama.cpp-turboquant-hip are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- World's first TurboQuant KV cache compression for llama.cpp on AMD ROCm (RX 9070 / gfx1201)☆19Apr 28, 2026Updated 3 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆23Jun 23, 2026Updated last month
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆16Apr 18, 2026Updated 3 months ago
- RDNA-native LLM inference engine in Rust.☆507Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆651Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- TurboQuant 3-bit KV-cache quantization for llama.cpp☆53Jul 30, 2026Updated 2 weeks ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆223Updated this week
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- LLM inference in C/C++☆2,262Updated this week
- ☆44Apr 26, 2026Updated 3 months ago
- ☆16Jan 20, 2026Updated 6 months ago
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆66Updated this week
- 🎭 Battle-tested plugins for Hermes Agent — zero core patches. Native vision bypass, multi-agent context injection, and more.☆47Jun 17, 2026Updated last month
- Evolution process to find the best quant tensor weights to build the most optimal GGUF options for an AI model.☆40May 13, 2026Updated 3 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆28Updated this week
- A wrapper that adds reveal effects, live ANSI recoloring, and keymaps to existing terminal apps.☆21Jun 8, 2026Updated 2 months ago
- ☆172Jun 14, 2026Updated last month
- The ultimate training toolkit for finetuning diffusion models. Add support for AMD ROCm GPUs repo.☆21Jan 29, 2026Updated 6 months ago
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆97May 23, 2026Updated 2 months ago
- Estimation and Prediction of Wet Season Calendar and Soil Water Balance for Agriculture☆16Feb 28, 2025Updated last year
- Effortlessly deploy a hardware-passthrough (VFIO) setup for Virtual Machines (VMs) on a Linux desktop. Includes quality-of-life features …☆18Apr 3, 2025Updated last year
- This is a repository for the Geospatial Data Abstraction Library (GDAL) and it's applications, examples and discussions in the world of s…☆10May 28, 2023Updated 3 years ago
- AI Model Demand Offloading Allocator☆53Aug 4, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Long Long Term Memory Neural Net Cells☆10Jan 25, 2022Updated 4 years ago
- Nacrith — Lossless text compression via ensemble neural arithmetic coding. Combines SmolLM2-135M language model with context mixing, adap…☆22Mar 21, 2026Updated 4 months ago
- The NowSecure Mobile Security Report☆11Nov 16, 2016Updated 9 years ago
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆263Updated this week
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22May 26, 2026Updated 2 months ago
- 🦉 OpenCode skill — error reflection & 5-Why root-cause learning agent☆18Jul 2, 2026Updated last month
- Fast color lookup for R☆10Feb 17, 2025Updated last year
- ☆25Jul 3, 2026Updated last month
- Create equal-area histograms with 'ggplot2'☆10Jul 23, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Rescue ramdisk for Pinephone written on Go☆10Apr 21, 2022Updated 4 years ago
- Obsidian plugin that surfaces semantically related notes using local AI embeddings. Privacy-first, offline-capable, 7 embedding providers…☆19May 22, 2026Updated 2 months ago
- Check Safety of SSH Public Keys☆12Oct 8, 2022Updated 3 years ago
- ☆11Feb 24, 2024Updated 2 years ago
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆16Apr 26, 2026Updated 3 months ago
- The Wetland DEM Ponding Model☆13Sep 22, 2021Updated 4 years ago
- Nested Dichotomy Logistic Regression Models☆10May 27, 2026Updated 2 months ago