NEW ROCmfp4 format for llama.cpp
☆157Jun 13, 2026Updated 3 months ago
Alternatives and similar repositories for rocmfp4-llama
Users that are interested in rocmfp4-llama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆409Sep 23, 2026Updated last week
- Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results…☆350Updated this week
- ☆56Sep 1, 2026Updated last month
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆76Sep 9, 2026Updated 3 weeks ago
- RDNA-native LLM inference engine in Rust.☆650Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 7…☆478Updated this week
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆18Apr 26, 2026Updated 5 months ago
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆29Jun 9, 2026Updated 3 months ago
- ☆1,972Updated this week
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆52May 10, 2026Updated 4 months ago
- ☆162Sep 2, 2026Updated last month
- ☆29Jul 30, 2026Updated 2 months ago
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 4 months ago
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆69Sep 5, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆764Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆73Aug 31, 2026Updated last month
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,143Updated this week
- ☆65Updated this week
- ☆109Aug 28, 2026Updated last month
- Linux driver for the embedded controller on the Sixunited AXB35-02 board.☆76Aug 20, 2026Updated last month
- A tiny, zero-dependency LLM agent harness that lives in your terminal. One static binary. No runtime, no daemon, no cloud. Local-LLM frie…☆46Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,887Updated this week
- ☆533Sep 22, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆87Updated this week
- Read-Only source code mirror, Proxmox uses mailing list workflow for development.☆27Aug 21, 2026Updated last month
- Official repository of AMD Playbooks☆110Updated this week
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆31Jul 7, 2026Updated 2 months ago
- Decentralizing distribution of open-source AI models.☆24Updated this week
- Technical docs to help you make you Halo Strix WORK!☆71Sep 3, 2026Updated last month
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,349Updated this week
- Qwen4-Exp (Qwen3.8-Flash-Next) SWA + MTP + PLE-on-iswa — bounded deep-context decode w/ long-range recall, TBQ4 4.25 bpv KV, DSpark + Rot…☆95Sep 24, 2026Updated last week
- OpenAI STT custom component for HA☆10Jan 25, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A PowerShell script to fully automate the setup of `llama.cpp` on Windows. It installs all prerequisites, including the correct CUDA Tool…☆24May 12, 2026Updated 4 months ago
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆81Sep 19, 2026Updated 2 weeks ago
- Show diff between 2 gem versions☆14Feb 12, 2019Updated 7 years ago
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 6 months ago
- ☆102Mar 8, 2026Updated 6 months ago
- ☆21Apr 14, 2026Updated 5 months ago
- ☆20Jul 21, 2026Updated 2 months ago