One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
☆227May 14, 2026Updated 2 months ago
Alternatives and similar repositories for qwen3.6-windows-server
Users that are interested in qwen3.6-windows-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆25May 8, 2026Updated 2 months ago
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆69Apr 29, 2026Updated 3 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆829Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆56Apr 28, 2026Updated 3 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,868Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)☆589Updated this week
- ☆118Apr 28, 2026Updated 3 months ago
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆91Updated this week
- Step-by-step guide: Qwen3.6 27B/35B on RTX PRO 6000 Blackwell — 120-200 tok/s with vLLM + MTP n=3☆16May 27, 2026Updated 2 months ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- Native Windows vLLM 0.26.0: CPython 3.13, CUDA 12.8, SM 7.5/8.6/8.9/12.0 for RTX 20/30/40/50, OpenAI-compatible serving, Triton/FlashAtte…☆36Updated this week
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆92Apr 28, 2026Updated 3 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆393Jun 16, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆2,990Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆167Updated this week
- Set your keyboard LEDs from the Mac OS X command-line (enables numlock support for CoolerMaster and other non-OSX keyboards)☆10Nov 23, 2019Updated 6 years ago
- ☆24Dec 29, 2025Updated 7 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 4 months ago
- Open-source framework for superagents.☆92Jun 16, 2026Updated last month
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆259Updated this week
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated last month
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆16May 11, 2026Updated 2 months ago
- Open source tool for transcirption and subtitling, alternative to happyscribe.☆36Feb 12, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆16Feb 9, 2026Updated 5 months ago
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆41Jul 2, 2026Updated last month
- ☆21Jun 13, 2024Updated 2 years ago
- Analysis of 44 AI agent frameworks through a context engineering lens. Feb 2026.☆16Feb 18, 2026Updated 5 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- LLM training & inference in python/C++ with web UI☆37Updated this week
- Vortex is a self-hosted RAG (Retrieval-Augmented Generation) application that lets you chat with your documents using any LLM provider. U…☆17Jul 11, 2026Updated 3 weeks ago
- FramePack with existing video input.☆29May 15, 2025Updated last year
- An MCP sse implementation of the Model Context Protocol (MCP) server integrated with SearXNG for providing AI agents with powerful, priva…☆24May 19, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A UWP application using the Lamp API to control the flashlight on a Windows device☆11Jun 17, 2023Updated 3 years ago
- Agent-agnostic tools for intelligent iterative optimization loops. Works with Claude Code, Codex, Cursor, OpenCode, Gemini CLI and more. …☆15Apr 2, 2026Updated 4 months ago
- Pixal3D image-to-3D nodes for ComfyUI - local TencentARC Pixal3D generation with textured GLB export + Windows support☆195Jun 12, 2026Updated last month
- ComfyUI custom_node to support both MoGe-2 and MoGe.☆22Mar 17, 2026Updated 4 months ago
- Some benchmark results of small models and quants that fit on DGX Spark☆49Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆431Jul 3, 2026Updated last month
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago