ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants
☆386Aug 22, 2026Updated 3 weeks ago
Alternatives and similar repositories for ROCmFPX
Users that are interested in ROCmFPX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NEW ROCmfp4 format for llama.cpp☆158Jun 13, 2026Updated 2 months ago
- Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audi…☆25Aug 16, 2026Updated 3 weeks ago
- RDNA-native LLM inference engine in Rust.☆625Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆73Aug 31, 2026Updated last week
- Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results…☆326Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆1,932Sep 5, 2026Updated last week
- Local AI setup for AMD Strix Halo APU - Lemonade + Vulkan + kyuz0☆73Updated this week
- Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI +…☆72Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆232Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆730Updated this week
- A high-performance RCCL / NCCL (ROCm Communication Collectives Library) plugin for Thunderbolt 5 that enables GPU-to-GPU communication ac…☆67Aug 14, 2026Updated 3 weeks ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,851Updated this week
- ☆108Aug 28, 2026Updated 2 weeks ago
- tested hermes agent recipes: configs, deploys, mcp, automations. copy, run, build.☆119Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A full GUI experience on top of llama.cpp: all-knobs model tuning, one-click build/update from upstream, HuggingFace discovery with VRAM-…☆81Updated this week
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated 2 months ago
- Collection of official scripts created by the Dione Team.☆15Feb 21, 2026Updated 6 months ago
- ☆518Sep 2, 2026Updated last week
- A tiny, zero-dependency LLM agent harness that lives in your terminal. One static binary. No runtime, no daemon, no cloud. Local-LLM frie…☆41Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,301Updated this week
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆67Sep 5, 2026Updated last week
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆164Aug 29, 2026Updated 2 weeks ago
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆17Apr 26, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated last month
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 3 months ago
- Train LLMs using a delightful web interface powered by magic ✨☆39Aug 28, 2026Updated 2 weeks ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,073Updated this week
- Unified DT for Xiaomi Redmi 4A/5A (rolex/riva)☆11Feb 27, 2026Updated 6 months ago
- Llama.cpp launcher with integrated huggingface☆71Updated this week
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆31Jul 7, 2026Updated 2 months ago
- Various scripts and tools to tinker with Huawei devices☆27May 11, 2026Updated 4 months ago
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆37Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆84Jul 12, 2026Updated 2 months ago
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆20Jul 19, 2026Updated last month
- DeepSeek 4 Flash local inference engine for Metal and CUDA with M5 optimizations.☆23May 24, 2026Updated 3 months ago
- Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.☆15Updated this week
- Read-Only source code mirror, Proxmox uses mailing list workflow for development.☆27Aug 21, 2026Updated 3 weeks ago
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆52May 10, 2026Updated 4 months ago
- A local-first Windows AI assistant with a multi-model council pipeline, Python sandbox, web search, and cloud AI support. Free. Private. …☆25Sep 1, 2026Updated last week