RDNA-native LLM inference engine in Rust.
☆489Jul 24, 2026Updated this week
Alternatives and similar repositories for hipfire
Users that are interested in hipfire are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆63Jul 10, 2026Updated 2 weeks ago
- NEW ROCmfp4 format for llama.cpp☆136Jun 13, 2026Updated last month
- ☆53Updated this week
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆136Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆624Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆62Jul 7, 2026Updated 2 weeks ago
- Fast LLM speculative inference server for consumer hardware.☆2,673Updated this week
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆16Apr 26, 2026Updated 2 months ago
- LLM inference in C/C++☆87Updated this week
- llama.cpp-gfx906☆139Mar 22, 2026Updated 4 months ago
- LLM inference in C/C++, but for GFX906!☆18Jul 15, 2026Updated last week
- ☆22May 12, 2026Updated 2 months ago
- ☆164Jun 14, 2026Updated last month
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,097Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,129Updated this week
- ☆1,779Jul 12, 2026Updated last week
- Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon☆487Updated this week
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆26Jun 9, 2026Updated last month
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,164Updated this week
- Random AI notes for working with local models or playing around with random machine learning bits.☆61Jun 7, 2026Updated last month
- TRELLIS.2 image-to-3D in C++/GGML (CUDA + Vulkan), with a resident HTTP server☆222Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆801Updated this week
- ☆472Jun 17, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.☆1,646Updated this week
- AI agent framework, written from scratch (not based on openclaw), focused on stripping it down to the bare necessities, optimizing token …☆413Updated this week
- ☆48Updated this week
- ☆19Nov 7, 2025Updated 8 months ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 2 months ago
- A harness optimized to smaller LLMs☆1,877Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆68Updated this week
- ☆16Updated this week
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆17Apr 18, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- llama.cpp fork with additional SOTA quants and improved performance☆2,960Updated this week
- Tensor library for machine learning☆35Updated this week
- Stop degrading your model's reasoning. A minimal, zero-config AI coding agent. Enforced ephemeral subagents keep context pure. From tiny …☆386Updated this week
- Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placem…☆257Updated this week
- ☆56Jun 8, 2026Updated last month
- ☆23Mar 21, 2026Updated 4 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week