A high-performance and light-weight router for vLLM large scale deployment
☆396Aug 31, 2026Updated last week
Alternatives and similar repositories for router
Users that are interested in router are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High-performance Rust benchmark client for vLLM serving endpoints.☆52Aug 3, 2026Updated last month
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆511Updated this week
- Agent skills for vLLM☆96Apr 3, 2026Updated 5 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,235Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 3 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- vLLM Daily Summarization of Merged PRs☆54Updated this week
- High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and S…☆201Updated this week
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,588Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆812Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,103Updated this week
- An LLM post-training framework with vLLM for RL Scaling☆454Updated this week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,516Updated this week
- Common recipes to run vLLM☆1,009Updated this week
- llm-d Router: The intelligent entry point for inference requests☆329Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,439Updated this week
- vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization☆2,560Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,986Updated this week
- A programmable Mixture-of-Models router for heterogeneous LLM inference☆5,640Updated this week
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆507Updated this week
- Efficient and easy multi-instance LLM serving☆562Mar 12, 2026Updated 5 months ago
- KV cache store for distributed LLM inference☆436Nov 13, 2025Updated 9 months ago
- Offline optimization of your disaggregated Dynamo graph☆437Updated this week
- The Intelligent Inference Scheduler for Large-scale Inference Services.☆69Feb 12, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,690Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,343Updated this week
- ☆385Jan 28, 2026Updated 7 months ago
- Distributed KV cache scheduling & offloading libraries☆176Updated this week
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆144Updated this week
- A PyTorch native library for training speculative decoding models☆243Updated this week
- Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serv…☆327Jul 29, 2026Updated last month
- ☆338Updated this week
- A framework for efficient model inference with omni-modality models☆6,702Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for…☆222Updated this week
- torchcomms: a modern PyTorch communications API☆391Updated this week
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Perplexity GPU Kernels☆607Nov 7, 2025Updated 10 months ago
- ☆16Apr 7, 2026Updated 5 months ago
- Community maintained hardware plugin for vLLM on Ascend☆2,768Updated this week
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,353Aug 23, 2026Updated 2 weeks ago