TurboQuant KV cache compression for llama.cpp — HIP/ROCm port for AMD RDNA3 (gfx1100)
☆55Apr 26, 2026Updated 2 months ago
Alternatives and similar repositories for llama.cpp-turboquant-hip
Users that are interested in llama.cpp-turboquant-hip are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- World's first TurboQuant KV cache compression for llama.cpp on AMD ROCm (RX 9070 / gfx1201)☆17Apr 28, 2026Updated 2 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆22Jun 23, 2026Updated 3 weeks ago
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆17Apr 18, 2026Updated 3 months ago
- RDNA-native LLM inference engine in Rust.☆486Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆623Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A way to visualize nginx config files☆10Nov 21, 2022Updated 3 years ago
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆29Jul 7, 2026Updated 2 weeks ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆221Jul 6, 2026Updated 2 weeks ago
- LLM inference in C/C++☆2,160Updated this week
- ☆44Apr 26, 2026Updated 2 months ago
- Rust libraries for reading Ultima Online data files☆15May 11, 2026Updated 2 months ago
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆62Jul 10, 2026Updated last week
- Historical PyTorch ROCm device-property patch for AMD RDNA GPUs. Kept for reference; not required to unlock RX 6000/7000 hardware.☆18Jun 20, 2026Updated last month
- Ultima Online T2A client recreated from Origin's 2.0.7 client decompilation☆15Jun 1, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Containerized Opinionated Agent Orchestration Platform☆15Mar 8, 2026Updated 4 months ago
- llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% t…☆307Updated this week
- 🎭 Battle-tested plugins for Hermes Agent — zero core patches. Native vision bypass, multi-agent context injection, and more.☆40Jun 17, 2026Updated last month
- Read-Only source code mirror, Proxmox uses mailing list workflow for development.☆26Jun 19, 2026Updated last month
- Desktop application for instant AI-powered text transformation. Translate, correct, summarize, and change the tone of any text, anywhere,…☆35Dec 29, 2025Updated 6 months ago
- Jailbreak PS4 FW9.00 for ESP8266. Compiles in the Arduino IDE.☆11Mar 18, 2022Updated 4 years ago
- A collection of LittleSnitch rule sets to block outbound connections to a given country.☆16Jun 13, 2023Updated 3 years ago
- ☆22Jul 12, 2026Updated last week
- Monorepo for sharing my most commonly used Nix expressions between projects.☆31Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆161Jun 14, 2026Updated last month
- Docker network setup with authelia, caddy, crowdsec, and wg-easy☆19Mar 14, 2026Updated 4 months ago
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆96May 23, 2026Updated last month
- Scripts and Ansible playbook to setup Arch Linux on ZFS.☆12Jun 16, 2024Updated 2 years ago
- Estimation and Prediction of Wet Season Calendar and Soil Water Balance for Agriculture☆16Feb 28, 2025Updated last year
- LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: sy …☆30May 13, 2026Updated 2 months ago
- This is a repository for the Geospatial Data Abstraction Library (GDAL) and it's applications, examples and discussions in the world of s…☆10May 28, 2023Updated 3 years ago
- Long Long Term Memory Neural Net Cells☆10Jan 25, 2022Updated 4 years ago
- The NowSecure Mobile Security Report☆11Nov 16, 2016Updated 9 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placem…☆256Updated this week
- Fast color lookup for R☆10Feb 17, 2025Updated last year
- Read Shapefiles☆11Mar 6, 2021Updated 5 years ago
- Create 'CREATE TABLE' and 'INSERT INTO...VALUES' SQL statements from R objects☆11Sep 11, 2022Updated 3 years ago
- Intra- and Inter-Regional Similarity☆13Feb 10, 2026Updated 5 months ago
- Create equal-area histograms with 'ggplot2'☆10Jul 23, 2023Updated 2 years ago
- Share your content of clipboard between 2 computers!☆10Apr 3, 2018Updated 8 years ago