vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
☆52May 10, 2026Updated 4 months ago
Alternatives and similar repositories for vllm-awq4-qwen
Users that are interested in vllm-awq4-qwen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆17Apr 26, 2026Updated 4 months ago
- Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results…☆326Updated this week
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆74Updated this week
- ☆518Sep 2, 2026Updated last week
- Experimental support for many TTS/STT LLMs wrapped in a Wyoming API for consumption via Homeassistant☆43Aug 25, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆73Aug 31, 2026Updated last week
- Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI +…☆72Updated this week
- Dockernized ComfyUI with PyTorch & flash-attention for gfx1151 (AMD Strix Halo, Ryzen AI Max+ 395), relying on AMD's pre-built and pre-co…☆45Feb 25, 2026Updated 6 months ago
- LLM inference in C/C++☆156Updated this week
- ☆20Sep 4, 2025Updated last year
- RDNA-native LLM inference engine in Rust.☆625Updated this week
- NEW ROCmfp4 format for llama.cpp☆158Jun 13, 2026Updated 2 months ago
- ☆1,932Sep 5, 2026Updated last week
- Run LLMs on AMD Ryzen AI NPU (Linux)☆29Mar 29, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LLM Fine Tuning Toolbox images for Ryzen AI 395+ Strix Halo☆65Sep 12, 2025Updated last year
- LLM inference in C/C++☆17Jul 29, 2026Updated last month
- ☆41Sep 22, 2025Updated 11 months ago
- Unified API for LLM providers and simple agent library in JS, Rust, and Go☆21Aug 3, 2026Updated last month
- Nemotron Speech ASR Docker deployment☆30Jun 7, 2026Updated 3 months ago
- World's first AMD NPU driver for TUXEDO laptops - Enable AI acceleration on Linux☆51Aug 5, 2025Updated last year
- Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.☆1,859Updated this week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆730Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,851Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆101Mar 8, 2026Updated 6 months ago
- ☆12Jul 3, 2020Updated 6 years ago
- Student little project to learn API uses and react☆13Nov 15, 2025Updated 9 months ago
- Mirror of my favourite hacking Zines for the lulz, nostalgy, and reference☆10Feb 18, 2020Updated 6 years ago
- Windows-native adaptation fork of Hermes Agent, based on upstream Hermes Agent 0.13.0. Improves local runtime environment, path handling,…☆20May 15, 2026Updated 3 months ago
- Charybdis Mini Travel Case☆17May 3, 2025Updated last year
- AMD iGPU AI Setup and Speed Test - GPD Pocket 4 - Linux + ROCm + Vulkan + AgentMake AI☆28Jun 17, 2026Updated 2 months ago
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- Build AI agents for your PC☆1,546Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Installer for the Better Minimap Cyberpunk 2077 mod☆11Mar 13, 2022Updated 4 years ago
- ☆175Apr 7, 2026Updated 5 months ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆17Updated this week
- Windows reference driver for Ether Dream☆11Apr 28, 2018Updated 8 years ago
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,696Updated this week
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆23Feb 18, 2026Updated 6 months ago
- hass☆10Jan 11, 2023Updated 3 years ago