Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.
☆64Jul 7, 2026Updated last month
Alternatives and similar repositories for kernel-anvil
Users that are interested in kernel-anvil are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM inference in C/C++, but for GFX906!☆23Updated this week
- llama.cpp-gfx906☆140Updated this week
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 2 months ago
- ☆18Jul 21, 2026Updated last month
- LLM inference in C/C++☆16Jul 29, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- RDNA-native LLM inference engine in Rust.☆546Updated this week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- Thin wrapper around GGML to make life easier☆48Jul 26, 2026Updated 3 weeks ago
- ☆63Jul 10, 2025Updated last year
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆21Feb 18, 2026Updated 6 months ago
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆60Aug 21, 2025Updated last year
- minimal fill-in-the-middle autocomplete for vscode/codium, for use with llama.cpp infill or any openai-compatible server☆23May 10, 2026Updated 3 months ago
- ☆62Updated this week
- Tensor library for machine learning☆36Jul 31, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆21Sep 4, 2025Updated 11 months ago
- LLM inference in C/C++☆18Updated this week
- ☆86Mar 23, 2026Updated 5 months ago
- Decentralizing distribution of open-source AI models.☆22Updated this week
- ☆25Feb 10, 2026Updated 6 months ago
- A PowerShell script to fully automate the setup of `llama.cpp` on Windows. It installs all prerequisites, including the correct CUDA Tool…☆17May 12, 2026Updated 3 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆46Updated this week
- Lightweight Python tool using Optuna for tuning llama.cpp flags: towards optimal tok/s for your machine☆53Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- ☆147Aug 11, 2026Updated last week
- ☆1,872Updated this week
- ☆250Oct 30, 2025Updated 9 months ago
- Enclosed chamber-type Delta 3d printer design☆11Aug 28, 2023Updated 2 years ago
- Extract a single expert from a Mixture Of Experts model using slerp interpolation.☆19May 26, 2024Updated 2 years ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆37Updated this week
- [ICASSP 2025] "FLowHigh: Towards efficient and high-quality audio super-resolution with single-step flow matching"☆34May 12, 2025Updated last year
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Local benchmarking UI for LLMs and AI agents☆23Apr 13, 2026Updated 4 months ago
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆319Updated this week
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆51May 10, 2026Updated 3 months ago
- llama.cpp fork with AMD XDNA2 NPU backend for Ryzen AI MAX (npu5/XDNA2) — matrix multiply offload via XRT☆20Apr 1, 2026Updated 4 months ago
- ☆44Apr 26, 2026Updated 3 months ago
- Llama.cpp launcher with integrated huggingface☆69Updated this week
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆99May 23, 2026Updated 3 months ago