One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
☆227May 14, 2026Updated 3 months ago
Alternatives and similar repositories for qwen3.6-windows-server
Users that are interested in qwen3.6-windows-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆26May 8, 2026Updated 3 months ago
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆73Apr 29, 2026Updated 3 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆906Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆57Apr 28, 2026Updated 3 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,068Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)☆602Jul 28, 2026Updated 3 weeks ago
- Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded …☆90Updated this week
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Aug 15, 2026Updated last week
- ☆25Apr 8, 2026Updated 4 months ago
- vLLM Docker Container for Qwen3.6 27b☆50Jun 15, 2026Updated 2 months ago
- ☆156Jun 13, 2026Updated 2 months ago
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆23Jul 2, 2026Updated last month
- 1CatV2 with TileLANG written FA-v100 and many goodies☆18Jul 15, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Hydro Desktop — A native desktop workspace for Claude Code, with project-based agents, Notebooks, capability management, scheduled tasks,…☆20Aug 15, 2026Updated last week
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆101Updated this week
- Test LLMs on real tasks. Compare models side-by-side.☆413Aug 10, 2026Updated last week
- llama.cpp fork with additional SOTA quants and improved performance☆3,077Aug 15, 2026Updated last week
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆170Updated this week
- Automatic benchmarking tool for all locally installed LM Studio models.☆19Aug 3, 2026Updated 3 weeks ago
- Set your keyboard LEDs from the Mac OS X command-line (enables numlock support for CoolerMaster and other non-OSX keyboards)☆10Nov 23, 2019Updated 6 years ago
- ☆24Dec 29, 2025Updated 7 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Open-source framework for superagents.☆94Aug 16, 2026Updated last week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆264Updated this week
- A repository to store helpful information and emerging insights in regard to LLMs☆21Oct 27, 2023Updated 2 years ago
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated 2 months ago
- This repo helps you understand how safetensors are structured to store different layers of an LLM and re-shard/re-chunk safetensors files…☆32Dec 14, 2025Updated 8 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆22Feb 9, 2026Updated 6 months ago
- Open source tool for transcirption and subtitling, alternative to happyscribe.☆36Feb 12, 2025Updated last year
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆48Jul 2, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Analysis of 44 AI agent frameworks through a context engineering lens. Feb 2026.☆17Feb 18, 2026Updated 6 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- Vortex is a self-hosted RAG (Retrieval-Augmented Generation) application that lets you chat with your documents using any LLM provider. U…☆17Jul 11, 2026Updated last month
- ☆23Jun 3, 2026Updated 2 months ago
- FramePack with existing video input.☆29May 15, 2025Updated last year
- 🔮 A powerful and stylish Prompt Generator powered by OpenAI and Python. Includes a built-in JSON editor, modular prompt libraries, and f…☆20Jul 12, 2025Updated last year
- XSLT procedural language for PostgreSQL☆20Mar 26, 2025Updated last year