One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
☆227May 14, 2026Updated 3 months ago
Alternatives and similar repositories for qwen3.6-windows-server
Users that are interested in qwen3.6-windows-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆27May 8, 2026Updated 4 months ago
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆74Apr 29, 2026Updated 4 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,073Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 4 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,228Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)☆627Aug 23, 2026Updated 2 weeks ago
- ☆119Apr 28, 2026Updated 4 months ago
- Qwen4-Exp (Qwen3.8-Flash-Next) SWA + MTP + PLE-on-iswa — bounded deep-context decode w/ long-range recall, TBQ4 4.25 bpv KV, DSpark + Rot…☆96Sep 2, 2026Updated last week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,851Updated this week
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Aug 15, 2026Updated 3 weeks ago
- ☆154Jun 13, 2026Updated 3 months ago
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- 1CatV2 with TileLANG written FA-v100 and many goodies☆19Jul 15, 2026Updated last month
- Hydro Desktop — A native desktop workspace for Claude Code, with project-based agents, Notebooks, capability management, scheduled tasks,…☆20Aug 15, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆107Aug 22, 2026Updated 3 weeks ago
- Test LLMs on real tasks. Compare models side-by-side.☆422Aug 10, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,217Updated this week
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆176Updated this week
- Set your keyboard LEDs from the Mac OS X command-line (enables numlock support for CoolerMaster and other non-OSX keyboards)☆10Nov 23, 2019Updated 6 years ago
- ☆24Dec 29, 2025Updated 8 months ago
- Open-source framework for superagents.☆106Updated this week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆271Updated this week
- Windows desktop control panel for local llama.cpp server☆380Jul 13, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated 2 months ago
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆25Feb 9, 2026Updated 7 months ago
- Open source tool for transcirption and subtitling, alternative to happyscribe.☆36Feb 12, 2025Updated last year
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆54Jul 2, 2026Updated 2 months ago
- ☆21Jun 13, 2024Updated 2 years ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- ☆23Jun 3, 2026Updated 3 months ago
- FramePack with existing video input.☆29May 15, 2025Updated last year
- A robust, single standalone html file AI interface for Ollama and OpenAI. Features local RAG with vector search, real-time voice/video ca…☆17Jul 4, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 🔮 A powerful and stylish Prompt Generator powered by OpenAI and Python. Includes a built-in JSON editor, modular prompt libraries, and f…☆20Jul 12, 2025Updated last year
- Native and Compact Structured Latents for 3D Generation☆214Apr 28, 2026Updated 4 months ago
- A UWP application using the Lamp API to control the flashlight on a Windows device☆11Jun 17, 2023Updated 3 years ago
- ☆22Apr 9, 2026Updated 5 months ago
- Pixal3D image-to-3D nodes for ComfyUI - local TencentARC Pixal3D generation with textured GLB export + Windows support☆210Jun 12, 2026Updated 3 months ago
- ComfyUI custom_node to support both MoGe-2 and MoGe.☆22Mar 17, 2026Updated 5 months ago
- Some benchmark results of small models and quants that fit on DGX Spark☆51Aug 23, 2026Updated 2 weeks ago