ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
☆347Sep 19, 2026Updated last week
Alternatives and similar repositories for ML-gfx906
Users that are interested in ML-gfx906 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆90Jun 23, 2026Updated 3 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 7 months ago
- llama.cpp-gfx906☆143Aug 23, 2026Updated last month
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆40Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆17Feb 21, 2026Updated 7 months ago
- Modified for AMD MI50 GPUs | A high-throughput and memory-efficient inference and serving engine for LLMs☆15Mar 2, 2026Updated 6 months ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆69Sep 5, 2026Updated 3 weeks ago
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆36Dec 15, 2025Updated 9 months ago
- Guidances for Test setup of 16 AMD MI50 32GB (for Deepseek v3.2)☆29May 11, 2026Updated 4 months ago
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆103Sep 8, 2026Updated 2 weeks ago
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,336Updated this week
- RDNA-native LLM inference engine in Rust.☆641Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The main repository for building Pascal-compatible versions of ML applications and libraries.☆218Aug 23, 2025Updated last year
- Random AI notes for working with local models or playing around with random machine learning bits.☆64Jun 7, 2026Updated 3 months ago
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆209Updated this week
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆26Jan 30, 2026Updated 7 months ago
- Torch-MIGraphX integrates AMD's graph inference engine with the PyTorch ecosystem.☆22Updated this week
- ☆13Dec 26, 2022Updated 3 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,260Updated this week
- Gives agents a real browser. URL in, pruned snapshot out. Replaces Playwright, Selenium, Puppeteer. Zero deps, zero wasted tokens.☆46Sep 20, 2026Updated last week
- ☆36Mar 26, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆461Updated this week
- A complete package that provides you with all the components needed to get started of dive deeper into Machine Learning Workloads on Cons…☆55Aug 24, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,883Updated this week
- AMD SMU Reverse Engineering (for AMD BC-250)☆20Dec 30, 2025Updated 8 months ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-reques…☆1,003Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,168Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- ☆445Apr 4, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆30Jul 24, 2026Updated 2 months ago
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆48Aug 17, 2026Updated last month
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆48Jan 13, 2026Updated 8 months ago
- A fork of vLLM enabling Pascal architecture GPUs☆37Feb 21, 2025Updated last year
- ☆17Oct 27, 2025Updated 11 months ago
- ☆29Jun 10, 2026Updated 3 months ago
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆404Updated this week