Incubating P/D sidecar for llm-d
☆17Nov 13, 2025Updated 8 months ago
Alternatives and similar repositories for llm-d-routing-sidecar
Users that are interested in llm-d-routing-sidecar are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Distributed KV cache scheduling & offloading libraries☆166Jul 28, 2026Updated last week
- llm-d benchmark scripts and tooling☆64Updated this week
- Simplified model deployment on llm-d☆29Jul 2, 2025Updated last year
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆17Updated this week
- llm-d Router: The intelligent entry point for inference requests☆278Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual h…☆175Updated this week
- Inference payload processor for llm-d☆16Updated this week
- ☆25Updated this week
- label ALL kubectl, kustomize, and helm objects, inline, without extra steps.(including namespaces and CRDs)☆15Apr 22, 2024Updated 2 years ago
- llm-d helm charts and deployment examples☆59May 1, 2026Updated 3 months ago
- Karmada功能、特性及源码分析☆16Mar 24, 2023Updated 3 years ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- d.run website☆18Jul 3, 2026Updated last month
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Kubernetes operator for bpfman☆39Jul 14, 2026Updated 2 weeks ago
- Automatically scales Kubernetes controllers to zero☆16May 30, 2019Updated 7 years ago
- Stateful API logic for agentic applications using vLLM☆59Updated this week
- Build and deploy Node.js application on Kubernetes☆16Sep 17, 2025Updated 10 months ago
- caniuse.com, but for kubernetes☆27Dec 25, 2024Updated last year
- 中国开发者活动日程(关注点:开源、开发者、云原生)☆27Updated this week
- ☆22Mar 11, 2026Updated 4 months ago
- Variant optimization autoscaler for distributed inference workloads☆52Updated this week
- GenAI inference performance benchmarking tool☆217Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Gateway API Inference Extension☆729Updated this week
- Ingress node firewall implements Kubernetes operator to provision stateless ingress node level firewall rules, stateless ingress node fir…☆76Updated this week
- Operator for managing Node Feature Discovery deployment☆78Mar 12, 2026Updated 4 months ago
- Let my Claude talk to yours.☆30Updated this week
- CLI for the Serverless Supercomputer☆25Sep 17, 2025Updated 10 months ago
- Operator for the mutating admission webhook for ClusterResourceOverride☆19Updated this week
- Prototypes and experiments for WG Device Management.☆16May 21, 2026Updated 2 months ago
- Kubernetes APIServer 高性能代理组件,代理 APIServer 的 List 请求,其它类型的请求会直接反向代理到原生 APIServer。 CKube 还额外支持了分页、搜索和索引等功能。 并且,CKube 100% 兼容原生 kubectl 和 ku…☆19Sep 16, 2022Updated 3 years ago
- helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/☆12Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆15Mar 6, 2025Updated last year
- 🧘 Extensive LLM endpoints, expended capabilities through your favorite protocols, 🕸️ GraphQL, ↔️ gRPC, ♾️ WebSocket. Extended SOTA supp…☆20Updated this week
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆483Updated this week
- ☆21Jul 20, 2026Updated 2 weeks ago
- NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across speci…☆40Updated this week
- Mix kubebuilder and code-generator example☆22Jul 3, 2020Updated 6 years ago
- Experimental DRA driver bringing CNI closer to Kubernetes☆44Oct 1, 2025Updated 10 months ago