RDNA-native LLM inference engine in Rust.
☆591Sep 2, 2026Updated this week
Alternatives and similar repositories for hipfire
Users that are interested in hipfire are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆70Updated this week
- NEW ROCmfp4 format for llama.cpp☆153Jun 13, 2026Updated 2 months ago
- ☆66Updated this week
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆370Aug 22, 2026Updated last week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆707Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆32Jul 7, 2026Updated last month
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆66Jul 7, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,828Updated this week
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆17Apr 26, 2026Updated 4 months ago
- LLM inference in C/C++☆154Updated this week
- Local LLM Server Manager + LlaMA.cpp + Chat☆130Updated this week
- ☆24May 12, 2026Updated 3 months ago
- ☆190Updated this week
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,579Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,553Updated this week
- ☆1,902Updated this week
- Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon☆509Updated this week
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆27Jun 9, 2026Updated 2 months ago
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,270Updated this week
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆99Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,016Updated this week
- ☆511Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.☆1,834Updated this week
- ☆82Updated this week
- AI agent framework, written from scratch (not based on openclaw), focused on stripping it down to the bare necessities, optimizing token …☆476Aug 18, 2026Updated 2 weeks ago
- ☆19Nov 7, 2025Updated 9 months ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 3 months ago
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆73Updated this week
- A harness optimized to smaller LLMs☆2,528Updated this week
- ☆20Jul 21, 2026Updated last month
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆16Apr 18, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- llama.cpp fork with additional SOTA quants and improved performance☆3,171Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆86Jun 23, 2026Updated 2 months ago
- Autonomous AI dev agent in pure Go. Enforced ephemeral subagents keep context pure. Making small local models viable for real dev work.☆415Updated this week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆268Updated this week
- ☆56Updated this week
- ☆25Jul 30, 2026Updated last month
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week