ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
☆335Aug 30, 2026Updated last week
Alternatives and similar repositories for ML-gfx906
Users that are interested in ML-gfx906 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆86Jun 23, 2026Updated 2 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 6 months ago
- llama.cpp-gfx906☆142Aug 23, 2026Updated 2 weeks ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆33Aug 25, 2026Updated last week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆16Feb 21, 2026Updated 6 months ago
- Modified for AMD MI50 GPUs | A high-throughput and memory-efficient inference and serving engine for LLMs☆15Mar 2, 2026Updated 6 months ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆67Updated this week
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆35Dec 15, 2025Updated 8 months ago
- triton3.2.0添加mi25/mi50/mi60支持☆14Apr 26, 2025Updated last year
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆101Updated this week
- The main repository for building Pascal-compatible versions of ML applications and libraries.☆218Aug 23, 2025Updated last year
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 5 months ago
- A llamacpp wrapper to manage and monitor your llama server instance over a web ui.☆28Jun 16, 2026Updated 2 months ago
- DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)☆26Aug 6, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,190Updated this week
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆452Sep 1, 2026Updated last week
- A complete package that provides you with all the components needed to get started of dive deeper into Machine Learning Workloads on Cons…☆54Aug 24, 2026Updated 2 weeks ago
- ☆15Jul 21, 2025Updated last year
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,839Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆925Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-reques…☆814Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- Flash Attention 2 implementation for Turing GPUs☆125Mar 23, 2026Updated 5 months ago
- unofficial support for rx580, rx470, Vega10, rx590, sp2048, or similar GFX803 GFX900 for zluda & Windows☆26Jan 17, 2026Updated 7 months ago
- ☆444Apr 4, 2025Updated last year
- ☆29Jul 24, 2026Updated last month
- High-performance FlashAttention-2 for AMD, Intel, and Apple GPUs. Drop-in replacement for PyTorch SDPA. Triton backend for ROCm (MI300X, …☆160Jan 27, 2026Updated 7 months ago
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆49Aug 17, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 7 months ago
- A fork of vLLM enabling Pascal architecture GPUs☆36Feb 21, 2025Updated last year
- ☆17Oct 27, 2025Updated 10 months ago
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆378Aug 22, 2026Updated 2 weeks ago
- SpartaDOS 3 compatible DOS for all Atari 8bit with at least 16k☆21May 3, 2026Updated 4 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,598Updated this week
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆210Updated this week