The Intelligent Inference Scheduler for Large-scale Inference Services.
☆69Feb 12, 2026Updated 6 months ago
Alternatives and similar repositories for aigw
Users that are interested in aigw are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- DeepXTrace is a lightweight tool for precisely diagnosing slow ranks in DeepEP-based environments.☆102Jan 16, 2026Updated 7 months ago
- PegaFlow is a high-performance KV cache offloading solution for vLLM v1 on single-node multi-GPU setups.☆25Jan 7, 2026Updated 7 months ago
- HTNN: A cloud-native gateway offering seamless extensibility for Istio and Envoy, in a native way by Go.☆124Updated this week
- A workload for deploying LLM inference services on Kubernetes☆283Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- LMCache-Ascend is a plugin for running LMCache on the Ascend NPU.☆85Updated this week
- 中国开发者活动日程(关注点:开源、开发者、云原生)☆29Updated this week
- A go sdk for coding Higress wasm plugins.☆32Aug 14, 2026Updated last week
- An LLM Mock Server that supports simulating the protocols of all LLM providers.☆16Jul 10, 2026Updated last month
- ☆325Updated this week
- What if everything is a io_uring?☆17Nov 10, 2022Updated 3 years ago
- ☆162Updated this week
- High Performance KV Cache Store for LLM☆59May 20, 2026Updated 3 months ago
- d.run website☆18Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Kubernetes-native AI serving platform for scalable model serving.☆432Updated this week
- Healthcheck library for OpenResty to validate upstream service status☆14Jul 21, 2026Updated last month
- ☆10Jun 27, 2024Updated 2 years ago
- 我陈平安,唯有一键,可搬山,倒海,降妖,镇魔,敕神,摘星,断江,摧城,开天!☆22Jun 4, 2022Updated 4 years ago
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆471Updated this week
- WebAssembly for Proxies (Go SDK)☆20May 25, 2026Updated 2 months ago
- Checkpoint-engine is a simple middleware to update model weights in LLM inference engines☆1,001Aug 12, 2026Updated last week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,351Updated this week
- Kubernetes CSI Driver for serving OCI model artifacts☆29May 25, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An OS kernel module for fast **remote** fork using advanced datacenter networking (RDMA).☆73Feb 15, 2025Updated last year
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆495Updated this week
- ☆36Jan 13, 2021Updated 5 years ago
- Provides deploy scripts and CSI for Lustre.☆14Apr 13, 2026Updated 4 months ago
- llm-d Router: The intelligent entry point for inference requests☆302Updated this week
- Offline optimization of your disaggregated Dynamo graph☆420Updated this week
- A userspace filesystem backing by Apache OpenDAL.☆47Updated this week
- Command-line tools for managing OCI model artifacts, which are bundled based on Model Spec☆79Aug 17, 2026Updated last week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- compatible library for ebpf programs to improve BTF portability☆14Jul 14, 2026Updated last month
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- An example of using hashicorp/raft☆18Dec 1, 2015Updated 10 years ago
- 简单的DNS代理,用来做内网DNS解析。☆20Nov 11, 2015Updated 10 years ago
- ☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!☆308Jan 26, 2026Updated 6 months ago
- Low level access to T-Head Xuantie RISC-V processors☆35Jul 26, 2026Updated 3 weeks ago
- Source for mosn.io site☆31Dec 3, 2025Updated 8 months ago