RDNA-native LLM inference engine in Rust.
☆511Aug 13, 2026Updated this week
Alternatives and similar repositories for hipfire
Users that are interested in hipfire are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆66Updated this week
- NEW ROCmfp4 format for llama.cpp☆144Jun 13, 2026Updated 2 months ago
- ☆56Updated this week
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆194Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆652Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆30Jul 7, 2026Updated last month
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated last month
- LLM speculative inference server for consumer & heterogeneous hardware☆2,739Updated this week
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆51May 10, 2026Updated 3 months ago
- LLM inference in C/C++☆136Updated this week
- llama.cpp-gfx906☆140Mar 22, 2026Updated 4 months ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆20Updated this week
- ☆23May 12, 2026Updated 3 months ago
- ☆173Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,347Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,345Updated this week
- ☆1,836Updated this week
- Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon☆498Jul 21, 2026Updated 3 weeks ago
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆26Jun 9, 2026Updated 2 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆62Jun 7, 2026Updated 2 months ago
- TRELLIS.2 image-to-3D in C++/GGML (CUDA + Vulkan), with a resident HTTP server☆248Aug 1, 2026Updated last week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆860Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆494Updated this week
- Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.☆1,730Updated this week
- ☆71Aug 1, 2026Updated last week
- AI agent framework, written from scratch (not based on openclaw), focused on stripping it down to the bare necessities, optimizing token …☆459Updated this week
- ☆19Nov 7, 2025Updated 9 months ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 2 months ago
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆70Updated this week
- A harness optimized to smaller LLMs☆2,373Jul 31, 2026Updated 2 weeks ago
- ☆16Jul 21, 2026Updated 3 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Frequency-based KV cache pruning for llama.cpp — 25% cache reduction, improved PPL at long context. GPU compaction kernel for HIP/ROCm.☆16Apr 18, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,028Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆81Jun 23, 2026Updated last month
- Stop degrading your model's reasoning. A minimal, zero-config AI coding agent. Enforced ephemeral subagents keep context pure. From tiny …☆402Updated this week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆264Updated this week
- ☆24Jul 30, 2026Updated 2 weeks ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week