Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.
☆67Sep 5, 2026Updated last week
Alternatives and similar repositories for kernel-anvil
Users that are interested in kernel-anvil are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- llama.cpp-gfx906☆143Aug 23, 2026Updated 3 weeks ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 3 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆87Jun 23, 2026Updated 2 months ago
- LLM inference in C/C++☆17Jul 29, 2026Updated last month
- RDNA-native LLM inference engine in Rust.☆625Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆338Aug 30, 2026Updated last week
- NEW ROCmfp4 format for llama.cpp☆158Jun 13, 2026Updated 2 months ago
- ☆518Sep 2, 2026Updated last week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 5 months ago
- Thin wrapper around GGML to make life easier☆48Jul 26, 2026Updated last month
- ☆63Jul 10, 2025Updated last year
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆23Feb 18, 2026Updated 6 months ago
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆61Aug 21, 2025Updated last year
- ☆73Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆20Sep 4, 2025Updated last year
- LLM inference in C/C++☆22Updated this week
- ☆87Mar 23, 2026Updated 5 months ago
- Decentralizing distribution of open-source AI models.☆24Updated this week
- ☆25Feb 10, 2026Updated 7 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- A PowerShell script to fully automate the setup of `llama.cpp` on Windows. It installs all prerequisites, including the correct CUDA Tool…☆21May 12, 2026Updated 4 months ago
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆74Updated this week
- Lightweight Python tool using Optuna for tuning llama.cpp flags: towards optimal tok/s for your machine☆59Aug 17, 2026Updated 3 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- High-performance late-interaction retrieval engine for on-prem AI. ColBERT/ColPali multi-vector search with Rust fused MaxSim, Triton GPU…☆17Jul 6, 2026Updated 2 months ago
- ☆1,932Sep 5, 2026Updated last week
- Enclosed chamber-type Delta 3d printer design☆11Aug 28, 2023Updated 3 years ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- ☆38Jun 10, 2026Updated 3 months ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆135Updated this week
- llama.cpp fork with AMD XDNA2 NPU backend for Ryzen AI MAX (npu5/XDNA2) — matrix multiply offload via XRT☆20Apr 1, 2026Updated 5 months ago
- ☆44Apr 26, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆52May 10, 2026Updated 4 months ago
- Llama.cpp launcher with integrated huggingface☆71Updated this week
- DEPRECATED. Moved to elm-community/json-extra =>☆10Jun 16, 2016Updated 10 years ago
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆386Aug 22, 2026Updated 3 weeks ago
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆102Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,301Updated this week
- ☆29Jan 19, 2026Updated 7 months ago