NEW ROCmfp4 format for llama.cpp
☆148Jun 13, 2026Updated 2 months ago
Alternatives and similar repositories for rocmfp4-llama
Users that are interested in rocmfp4-llama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆319Updated this week
- Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and…☆292Updated this week
- ☆56Jul 31, 2026Updated 3 weeks ago
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆69Updated this week
- RDNA-native LLM inference engine in Rust.☆546Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆17Apr 26, 2026Updated 3 months ago
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆26Jun 9, 2026Updated 2 months ago
- ☆1,872Updated this week
- Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI +…☆68Updated this week
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆51May 10, 2026Updated 3 months ago
- ☆147Aug 11, 2026Updated last week
- ☆25Jul 30, 2026Updated 3 weeks ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 3 months ago
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆685Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆70Updated this week
- ☆79Aug 1, 2026Updated 3 weeks ago
- ☆41Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆906Updated this week
- ☆100Aug 10, 2026Updated last week
- LLM inference in C/C++☆144Updated this week
- Linux driver for the embedded controller on the Sixunited AXB35-02 board.☆62Updated this week
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A tiny, zero-dependency LLM agent harness that lives in your terminal. One static binary. No runtime, no daemon, no cloud. Local-LLM frie…☆38Aug 9, 2026Updated 2 weeks ago
- ☆505Updated this week
- Official repository of AMD Playbooks☆98Updated this week
- Technical docs to help you make you Halo Strix WORK!☆71Jan 10, 2026Updated 7 months ago
- Experimental llama.cpp fork for inference research and development☆779Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,228Updated this week
- Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded …☆90Updated this week
- OpenAI STT custom component for HA☆11Jan 25, 2024Updated 2 years ago
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated last month
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 5 months ago
- ☆101Mar 8, 2026Updated 5 months ago
- ☆17Apr 14, 2026Updated 4 months ago
- ☆18Jul 21, 2026Updated last month
- ☆21Mar 28, 2026Updated 4 months ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆35Updated this week
- Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.☆45Aug 16, 2026Updated last week