RDNA-native LLM inference engine in Rust.
☆635Sep 21, 2026Updated this week
Alternatives and similar repositories for hipfire
Users that are interested in hipfire are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆76Sep 9, 2026Updated last week
- NEW ROCmfp4 format for llama.cpp☆158Jun 13, 2026Updated 3 months ago
- ☆80Updated this week
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆397Aug 22, 2026Updated last month
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆745Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆52May 10, 2026Updated 4 months ago
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆31Jul 7, 2026Updated 2 months ago
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆69Sep 5, 2026Updated 2 weeks ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,872Updated this week
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆18Apr 26, 2026Updated 4 months ago
- llama.cpp-gfx906☆143Aug 23, 2026Updated last month
- ☆27May 12, 2026Updated 4 months ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆140Sep 7, 2026Updated 2 weeks ago
- ☆199Sep 2, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Sep 13, 2026Updated last week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,767Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,722Updated this week
- ☆1,949Updated this week
- Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon☆513Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,329Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,109Updated this week
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆103Sep 8, 2026Updated 2 weeks ago
- ☆524Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.☆1,891Updated this week
- ☆86Updated this week
- handwritten harness for AI, not vibecoded, everything is a module/plugin, tailor made for use with local models, tiny system prompt (arou…☆492Updated this week
- ☆19Nov 7, 2025Updated 10 months ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 4 months ago
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆73Aug 31, 2026Updated 3 weeks ago
- ☆20Jul 21, 2026Updated 2 months ago
- A harness optimized to smaller LLMs☆2,611Updated this week
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆16Apr 18, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- llama.cpp fork with additional SOTA quants and improved performance☆3,254Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆87Jun 23, 2026Updated 3 months ago
- Tensor library for machine learning☆36Jul 31, 2026Updated last month
- High-performance AI agent for long-horizon tasks. Built on empirical research. 200k+ tokens of work inside a 64k context window.☆431Updated this week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆275Updated this week
- ☆56Sep 1, 2026Updated 3 weeks ago
- ☆28Jul 30, 2026Updated last month