A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels)
☆592Jul 28, 2026Updated 2 weeks ago
Alternatives and similar repositories for vllm-windows
Users that are interested in vllm-windows are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆16Jul 1, 2026Updated last month
- Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on t…☆25May 8, 2026Updated 3 months ago
- Fork of the Triton language and compiler for Windows support and easy installation☆1,958Feb 18, 2026Updated 5 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆227May 14, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,026Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,127Updated this week
- Triton with Windows support☆227Updated this week
- Pynini for Windows☆25Apr 11, 2025Updated last year
- Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossi…☆154Jul 9, 2026Updated last month
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆852Updated this week
- This is a pre-built wheel of Triton 3.3.0 for Windows with Nvidia only + Proton☆44May 18, 2025Updated last year
- Deepspeed windows information☆44Mar 9, 2024Updated 2 years ago
- Pre-compiled Python whl for Flash-attention, SageAttention, NATTEN, xFormer etc☆797Jul 31, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Run GGUF models easily with a KoboldAI UI. One File. Zero Install.☆11,388Updated this week
- Official ConvRot implementation. A plug-and-play, convolution-like rotation module enabling efficient W4A4 quantization for diffusion mod…☆23Jul 3, 2026Updated last month
- Fast and memory-efficient exact attention☆944Dec 9, 2025Updated 8 months ago
- Flash Attention WHL Builds for Windows☆16May 30, 2025Updated last year
- A graphical tool to simplify building and installing DeepSpeed 0.15.x or later on Windows systems.☆21Apr 26, 2025Updated last year
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆167Updated this week
- [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models☆3,931Mar 7, 2026Updated 5 months ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆88,737Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,323Updated this week
- Advanced CLI diffusion inference/training suite based on Musubi Tuner☆40Apr 15, 2026Updated 3 months ago
- A transformers implementation of csm-streaming☆30May 16, 2025Updated last year
- C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal for…☆531Updated this week
- LLM inference in C/C++☆2,258Updated this week
- Omni inference in C/C++☆241Updated this week
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆124Jul 22, 2026Updated 2 weeks ago
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆76Jan 19, 2026Updated 6 months ago
- GGUF Quantization support for native ComfyUI models☆3,895Jan 12, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LLM inference in C/C++☆123,460Updated this week
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- gguf (GPT-Generated Unified Format) connector☆58Updated this week
- Experimental llama.cpp fork for inference research and development☆735Updated this week
- Fork of SageAttention for Windows wheels and easy installation☆968Jul 17, 2026Updated 3 weeks ago
- ☆168Updated this week
- miaoshouai-assistant for webui-forge☆15Aug 15, 2024Updated last year