A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)
☆577Jul 12, 2026Updated last week
Alternatives and similar repositories for vllm-windows
Users that are interested in vllm-windows are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆15Jul 1, 2026Updated 3 weeks ago
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆25May 8, 2026Updated 2 months ago
- Fork of the Triton language and compiler for Windows support and easy installation☆1,955Feb 18, 2026Updated 5 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆223May 14, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,950Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,061Updated this week
- Triton with Windows support☆192Updated this week
- Pynini for Windows☆24Apr 11, 2025Updated last year
- Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossi…☆151Jul 9, 2026Updated last week
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆797Updated this week
- This is a pre-built wheel of Triton 3.3.0 for Windows with Nvidia only + Proton☆44May 18, 2025Updated last year
- Deepspeed windows information☆44Mar 9, 2024Updated 2 years ago
- Pre-compiled Python whl for Flash-attention, SageAttention, NATTEN, xFormer etc☆719Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Run GGUF models easily with a KoboldAI UI. One File. Zero Install.☆11,089Updated this week
- Official ConvRot implementation. A plug-and-play, convolution-like rotation module enabling efficient W4A4 quantization for diffusion mod…☆20Jul 3, 2026Updated 2 weeks ago
- Fast and memory-efficient exact attention☆943Dec 9, 2025Updated 7 months ago
- Flash Attention WHL Builds for Windows☆16May 30, 2025Updated last year
- A graphical tool to simplify building and installing DeepSpeed 0.15.x or later on Windows systems.☆21Apr 26, 2025Updated last year
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆165Updated this week
- [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models☆3,911Mar 7, 2026Updated 4 months ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆86,804Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,103Updated this week
- Advanced CLI diffusion inference/training suite based on Musubi Tuner☆40Apr 15, 2026Updated 3 months ago
- Lower Precision Floating Point Operations☆83Feb 22, 2026Updated 5 months ago
- A transformers implementation of csm-streaming☆30May 16, 2025Updated last year
- C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal for…☆469Updated this week
- LLM inference in C/C++☆2,164Updated this week
- Omni inference in C/C++☆214Updated this week
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆123Updated this week
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆69Jan 19, 2026Updated 6 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- GGUF Quantization support for native ComfyUI models☆3,840Jan 12, 2026Updated 6 months ago
- LLM inference in C/C++☆121,220Updated this week
- Fast LLM speculative inference server for consumer hardware.☆2,666Updated this week
- gguf (GPT-Generated Unified Format) connector☆57Updated this week
- LLAMA Turboquant implementation with CUDA support☆705Updated this week
- Fork of SageAttention for Windows wheels and easy installation☆885Updated this week
- ☆144Updated this week