vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
☆51May 10, 2026Updated 3 months ago
Alternatives and similar repositories for vllm-awq4-qwen
Users that are interested in vllm-awq4-qwen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆17Apr 26, 2026Updated 3 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 2 months ago
- Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and…☆292Updated this week
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆69Updated this week
- ☆505Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below☆70Updated this week
- ☆62Updated this week
- ☆28Jun 10, 2026Updated 2 months ago
- Dockernized ComfyUI with PyTorch & flash-attention for gfx1151 (AMD Strix Halo, Ryzen AI Max+ 395), relying on AMD's pre-built and pre-co…☆46Feb 25, 2026Updated 5 months ago
- LLM inference in C/C++☆144Updated this week
- NEW ROCmfp4 format for llama.cpp☆148Jun 13, 2026Updated 2 months ago
- RDNA-native LLM inference engine in Rust.☆546Updated this week
- ☆1,872Updated this week
- LLM Fine Tuning Toolbox images for Ryzen AI 395+ Strix Halo☆66Sep 12, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Linux driver for the embedded controller on the Sixunited AXB35-02 board.☆62Updated this week
- LLM inference in C/C++☆16Jul 29, 2026Updated 3 weeks ago
- DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)☆25Aug 6, 2026Updated 2 weeks ago
- A high-performance RCCL / NCCL (ROCm Communication Collectives Library) plugin for Thunderbolt 5 that enables GPU-to-GPU communication ac…☆61Aug 14, 2026Updated last week
- Fresh builds of llama.cpp with AMD ROCm™ 7 acceleration☆685Updated this week
- Local AI setup for AMD Strix Halo APU - Lemonade + Vulkan + kyuz0☆71Updated this week
- MCP server for creating UI flowcharts☆11Jan 5, 2025Updated last year
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.☆1,791Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆177Aug 9, 2026Updated 2 weeks ago
- Experimental implementation of DeepSeek v4 flaash in llama.cpp☆24Apr 30, 2026Updated 3 months ago
- Windows-native adaptation fork of Hermes Agent, based on upstream Hermes Agent 0.13.0. Improves local runtime environment, path handling,…☆20May 15, 2026Updated 3 months ago
- ☆21Sep 4, 2025Updated 11 months ago
- Automatically identify, retrieve and store memories from user conversations in Open WebUI.☆19Sep 25, 2025Updated 10 months ago
- Easy access to unsaved files for vscode.☆12Jul 10, 2025Updated last year
- Run llama.cpp server with Vulkan☆19Aug 5, 2026Updated 2 weeks ago
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆30Jul 7, 2026Updated last month
- Mistral Vibe rewritten in Rust by Devstral 2☆21Dec 23, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆23Jul 2, 2026Updated last month
- ☆44Apr 26, 2026Updated 3 months ago
- pi.dev + llama.cpp ❤️☆24Jul 6, 2026Updated last month
- A simple Pong game in Flutter☆10Oct 19, 2019Updated 6 years ago
- ☆15Feb 10, 2020Updated 6 years ago
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆21Feb 18, 2026Updated 6 months ago
- hass☆10Jan 11, 2023Updated 3 years ago