vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
☆18Apr 26, 2026Updated 5 months ago
Alternatives and similar repositories for vllm-qwen
Users that are interested in vllm-qwen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆53May 10, 2026Updated 4 months ago
- ☆87Updated this week
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆29Jun 9, 2026Updated 3 months ago
- ☆56Sep 1, 2026Updated last month
- Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI +…☆76Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Portable vLLM builds with AMD ROCm acceleration for Lemonade☆75Sep 9, 2026Updated 3 weeks ago
- ☆29Jul 30, 2026Updated 2 months ago
- ☆20Jul 21, 2026Updated 2 months ago
- Local AI setup for AMD Strix Halo APU - Lemonade + Vulkan + kyuz0☆84Sep 16, 2026Updated 2 weeks ago
- Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results…☆350Updated this week
- Ameoba Single key PCB variant designed for simplicity☆13Apr 17, 2023Updated 3 years ago
- NEW ROCmfp4 format for llama.cpp☆157Jun 13, 2026Updated 3 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆64Jun 7, 2026Updated 3 months ago
- ☆13Nov 19, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Trackpoint emulator: 3 choc swithes and analog stick☆15Mar 23, 2025Updated last year
- Constructive Solid Geometry (CSG) and Quick Hull Library for golang☆10Oct 30, 2019Updated 6 years ago
- Yet NOT Another Android Syntax Highlighter (YNAASH)☆15Jun 19, 2026Updated 3 months ago
- Script to normalize and validate STL files so that they play better with git version control.☆15Mar 15, 2021Updated 5 years ago
- ☆16Feb 1, 2021Updated 5 years ago
- Decentralizing distribution of open-source AI models.☆25Updated this week
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- MCP that makes Xcode 26.3's MCP compatible with Cursor and other strict MCP-spec-compliant clients☆20Sep 25, 2026Updated last week
- Rest AIP server for commuticate with Tion devices☆12Dec 19, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A simple but instructive implementation of DP, TP, FSDP, FSDP+TP using pytorch distributed primitives☆24Apr 12, 2026Updated 5 months ago
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆19Jul 13, 2026Updated 2 months ago
- [Archived] No longer maintained. Run AI coding agents in parallel inside hardened, isolated containers.☆27Updated this week
- Parametric 3D CAD with a real B-rep kernel (OpenCASCADE)☆445Updated this week
- Realistic email testing server in a single Docker container. SMTP, IMAP, DKIM/DMARC, spam filtering, REST API, Node/Python SDKs, and an M…☆23Updated this week
- 🦉 OpenCode skill — error reflection & 5-Why root-cause learning agent☆17Sep 1, 2026Updated last month
- 24 keys ought to be enough for anybody☆15May 28, 2024Updated 2 years ago
- CEREBRO-RED v2: Advanced LLM Red Team Research Platform with PAIR Algorithm and LLM-as-a-Judge Evaluation☆16Mar 21, 2026Updated 6 months ago
- An Intellij IDEA and Android Studio Plugin to de-obfuscate your Proguard stack traces☆16Sep 2, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Share your content of clipboard between 2 computers!☆10Apr 3, 2018Updated 8 years ago
- ☆18Feb 2, 2022Updated 4 years ago
- A Kanban board-like UI to manage agentic workflows. Designed for seamless integration with KaibanJS to visualize and manage AI agent task…☆21May 15, 2026Updated 4 months ago
- Deterministic Linux runtime enforcement with eBPF LSM: block file/network operations before syscalls complete.☆36Updated this week
- OccJava - A SWIG-generated Java wrapper for OpenCascade☆20Mar 22, 2016Updated 10 years ago
- ☆11Jun 21, 2023Updated 3 years ago
- RDNA-native LLM inference engine in Rust.☆652Updated this week