NEW ROCmfp4 format for llama.cpp
☆139Jun 13, 2026Updated last month
Alternatives and similar repositories for rocmfp4-llama
Users that are interested in rocmfp4-llama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆162Updated this week
- AMD Strix Halo / Ryzen AI Halo local LLM setup and benchmark guide for Ryzen AI MAX+ 395 and Radeon 8060S: Ollama, llama.cpp Vulkan/RADV,…☆255Updated this week
- ☆56Updated this week
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆65Jul 10, 2026Updated 3 weeks ago
- RDNA-native LLM inference engine in Rust.☆496Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆16Apr 26, 2026Updated 3 months ago
- ☆1,802Updated this week
- Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI +…☆62Updated this week
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆50May 10, 2026Updated 2 months ago
- ☆145May 29, 2026Updated 2 months ago
- ☆24Updated this week
- Curated AMD GPU compatibility index and CLI for AI workloads☆31May 23, 2026Updated 2 months ago
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆636Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆69Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆57Updated this week
- ☆25Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆829Updated this week
- LLM inference in C/C++☆124Updated this week
- Linux driver for the embedded controller on the Sixunited AXB35-02 board.☆59Jul 11, 2026Updated 3 weeks ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- A tiny, zero-dependency LLM agent harness that lives in your terminal. One static binary. No runtime, no daemon, no cloud. Local-LLM frie…☆33Updated this week
- ☆55Updated this week
- ☆487Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Read-Only source code mirror, Proxmox uses mailing list workflow for development.☆27Jun 19, 2026Updated last month
- Official repository of AMD Playbooks☆87Updated this week
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆30Jul 7, 2026Updated 3 weeks ago
- Technical docs to help you make you Halo Strix WORK!☆71Jan 10, 2026Updated 6 months ago
- Experimental llama.cpp fork for inference research and development☆725Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,183Updated this week
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆91Updated this week
- A PowerShell script to fully automate the setup of `llama.cpp` on Windows. It installs all prerequisites, including the correct CUDA Tool…☆16May 12, 2026Updated 2 months ago
- OpenAI STT custom component for HA☆11Jan 25, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆77Jul 14, 2026Updated 3 weeks ago
- Show diff between 2 gem versions☆14Feb 12, 2019Updated 7 years ago
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 4 months ago
- ☆98Mar 8, 2026Updated 4 months ago
- ☆16Apr 14, 2026Updated 3 months ago
- Tensor library for machine learning☆36Updated this week
- ☆21Mar 28, 2026Updated 4 months ago