A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)
☆645Sep 18, 2026Updated this week
Alternatives and similar repositories for vllm-windows
Users that are interested in vllm-windows are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆19Jul 1, 2026Updated 2 months ago
- Native Windows vLLM: stable 0.27.1; 0.29.0 prereleases for Python 3.13/CUDA 13.0 and Python 3.14/CUDA 13.2 (CPU TorchAudio). PyTorch 2.13…☆58Updated this week
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆28May 8, 2026Updated 4 months ago
- Fork of the Triton language and compiler for Windows support and easy installation☆1,954Feb 18, 2026Updated 7 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆229May 14, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- llama.cpp fork with additional SOTA quants and improved performance☆3,251Updated this week
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,472Updated this week
- Triton with Windows support☆257Updated this week
- Pynini for Windows☆26Apr 11, 2025Updated last year
- Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossi…☆157Jul 9, 2026Updated 2 months ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Sep 13, 2026Updated last week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,105Updated this week
- This is a pre-built wheel of Triton 3.3.0 for Windows with Nvidia only + Proton☆44May 18, 2025Updated last year
- Deepspeed windows information☆44Mar 9, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Pre-compiled Python whl for Flash-attention, SageAttention, NATTEN, xFormer etc☆867Sep 7, 2026Updated 2 weeks ago
- Run GGUF models easily with a KoboldAI UI. One File. Zero Install.☆11,812Updated this week
- Official ConvRot implementation. A plug-and-play, convolution-like rotation module enabling efficient W4A4 quantization for diffusion mod…☆32Jul 3, 2026Updated 2 months ago
- Fast and memory-efficient exact attention☆944Dec 9, 2025Updated 9 months ago
- Flash Attention WHL Builds for Windows☆16May 30, 2025Updated last year
- A graphical tool to simplify building and installing DeepSpeed 0.15.x or later on Windows systems.☆21Apr 26, 2025Updated last year
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆177Updated this week
- [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models☆3,951Sep 6, 2026Updated 2 weeks ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,709Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A high-throughput and memory-efficient inference and serving engine for LLMs☆92,343Updated this week
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- Advanced CLI diffusion inference/training suite based on Musubi Tuner☆39Apr 15, 2026Updated 5 months ago
- Lower Precision Floating Point Operations☆86Feb 22, 2026Updated 6 months ago
- A transformers implementation of csm-streaming☆30May 16, 2025Updated last year
- LLM inference in C/C++☆2,395Updated this week
- C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal for…☆656Updated this week
- GGUF Quantization support for native ComfyUI models☆4,041Jan 12, 2026Updated 8 months ago
- Omni inference in C/C++☆267Sep 11, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆127Jul 22, 2026Updated last month
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆80Jan 19, 2026Updated 8 months ago
- LLM inference in C/C++☆129,071Updated this week
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆167Sep 11, 2026Updated last week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,870Updated this week
- gguf (GPT-Generated Unified Format) connector☆61Sep 14, 2026Updated last week
- Experimental llama.cpp fork for inference research and development☆870Updated this week