vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
☆50May 10, 2026Updated 2 months ago
Alternatives and similar repositories for vllm-awq4-qwen
Users that are interested in vllm-awq4-qwen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Random AI notes for working with local models or playing around with random machine learning bits.☆61Jun 7, 2026Updated last month
- ☆17Jul 21, 2026Updated last week
- AMD Strix Halo / Ryzen AI Halo local LLM setup and benchmark guide for Ryzen AI MAX+ 395 and Radeon 8060S: Ollama, llama.cpp Vulkan/RADV,…☆255Updated this week
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆65Jul 10, 2026Updated 3 weeks ago
- ☆487Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Experimental support for many TTS/STT LLMs wrapped in a Wyoming API for consumption via Homeassistant☆37Jun 21, 2026Updated last month
- ☆24Updated this week
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆69Updated this week
- ☆55Updated this week
- LLM inference in C/C++☆124Updated this week
- ☆145May 29, 2026Updated 2 months ago
- ☆58Updated this week
- RDNA-native LLM inference engine in Rust.☆496Updated this week
- ☆1,802Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Local inference with GMKTek evo-x2☆26Nov 3, 2025Updated 9 months ago
- LLM Fine Tuning Toolbox images for Ryzen AI 395+ Strix Halo☆64Sep 12, 2025Updated 10 months ago
- Linux driver for the embedded controller on the Sixunited AXB35-02 board.☆59Jul 11, 2026Updated 3 weeks ago
- LLM inference in C/C++☆16Updated this week
- DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)☆19Jul 19, 2026Updated 2 weeks ago
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆636Updated this week
- Local AI setup for AMD Strix Halo APU - Lemonade + Vulkan + kyuz0☆62Updated this week
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- ☆250Oct 30, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.☆1,662Updated this week
- ☆21Sep 4, 2025Updated 11 months ago
- Run llama.cpp server with Vulkan☆19Jul 26, 2026Updated last week
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆30Jul 7, 2026Updated 3 weeks ago
- A converter for transferring gguf Q4_0, Q4_1 to FLM Q4NX☆37May 27, 2026Updated 2 months ago
- Installing NFS on a Buffalo 220 NAS device☆10Nov 27, 2019Updated 6 years ago
- pi.dev + llama.cpp ❤️☆25Jul 6, 2026Updated 3 weeks ago
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆20Feb 18, 2026Updated 5 months ago
- Windows reference driver for Ether Dream☆11Apr 28, 2018Updated 8 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated 3 weeks ago
- hass☆10Jan 11, 2023Updated 3 years ago
- Create 3D files in the CLI with Small Language Model☆44Oct 15, 2025Updated 9 months ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago
- GT Bike V - Docs☆14Jan 25, 2023Updated 3 years ago
- Technical docs to help you make you Halo Strix WORK!☆71Jan 10, 2026Updated 6 months ago
- Your AI's anchor to reality. ⚓☆61Jul 9, 2026Updated 3 weeks ago