One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
☆230May 14, 2026Updated 4 months ago
Alternatives and similar repositories for qwen3.6-windows-server
Users that are interested in qwen3.6-windows-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆28May 8, 2026Updated 4 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,143Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 5 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,333Updated this week
- ☆118Apr 28, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)☆655Sep 18, 2026Updated 2 weeks ago
- Qwen4-Exp (Qwen3.8-Flash-Next) SWA + MTP + PLE-on-iswa — bounded deep-context decode w/ long-range recall, TBQ4 4.25 bpv KV, DSpark + Rot…☆95Sep 24, 2026Updated last week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,887Updated this week
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Updated this week
- ☆25Apr 8, 2026Updated 5 months ago
- vLLM Docker Container for the latest Qwen 27b☆52Updated this week
- ☆152Jun 13, 2026Updated 3 months ago
- Native Windows vLLM: 0.27.1 Latest; 0.29.0 Python 3.14/CUDA 13.2 release (CPU TorchAudio), plus Python 3.13/CUDA 13.0 prerelease. PyTorch…☆59Sep 22, 2026Updated last week
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆112Aug 22, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Test LLMs on real tasks. Compare models side-by-side.☆429Aug 10, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,272Updated this week
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆180Updated this week
- Automatic benchmarking tool for all locally installed LM Studio models.☆20Sep 4, 2026Updated 3 weeks ago
- Set your keyboard LEDs from the Mac OS X command-line (enables numlock support for CoolerMaster and other non-OSX keyboards)☆10Nov 23, 2019Updated 6 years ago
- ☆24Dec 29, 2025Updated 9 months ago
- Open-source framework for superagents.☆114Sep 11, 2026Updated 3 weeks ago
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆280Updated this week
- Windows desktop control panel for local llama.cpp server☆385Jul 13, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated 3 months ago
- Open source tool for transcirption and subtitling, alternative to happyscribe.☆37Feb 12, 2025Updated last year
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆54Jul 2, 2026Updated 3 months ago
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆29Feb 9, 2026Updated 7 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆133Updated this week
- Vortex is a self-hosted RAG (Retrieval-Augmented Generation) application that lets you chat with your documents using any LLM provider. U…☆17Sep 3, 2026Updated last month
- ☆25Jun 3, 2026Updated 4 months ago
- XSLT procedural language for PostgreSQL☆20Mar 26, 2025Updated last year
- An MCP sse implementation of the Model Context Protocol (MCP) server integrated with SearXNG for providing AI agents with powerful, priva…☆27May 19, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Native and Compact Structured Latents for 3D Generation☆218Apr 28, 2026Updated 5 months ago
- Agent-agnostic tools for intelligent iterative optimization loops. Works with Claude Code, Codex, Cursor, OpenCode, Gemini CLI and more. …☆18Apr 2, 2026Updated 6 months ago
- A fast and flexible random prompt generator for ComfyUI with 12 columns (Empty / Pre-filled SFW / Pre-filled NSFW)☆21Dec 8, 2025Updated 9 months ago
- MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)☆211Sep 21, 2026Updated last week
- This repo contains the critical chat_template.jinja fix for Qwen3/3.5/3.6 to work smoothly in agentic task.☆96May 2, 2026Updated 5 months ago
- TRELLIS.2 image-to-3D in C++/GGML (CUDA + Vulkan), with a resident HTTP server☆329Sep 25, 2026Updated last week
- LLM inference in C/C++☆2,414Updated this week