llm-d benchmark scripts and tooling
☆67Sep 4, 2026Updated this week
Alternatives and similar repositories for llm-d-benchmark
Users that are interested in llm-d-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Simplified model deployment on llm-d☆28Jul 2, 2025Updated last year
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 9 months ago
- GenAI inference performance benchmarking tool☆238Updated this week
- Distributed KV cache scheduling & offloading libraries☆176Updated this week
- Helm charts for llm-d☆52Jul 22, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- llm-d Router: The intelligent entry point for inference requests☆329Updated this week
- helm charts for deploying models with llm-d☆32Updated this week
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual h…☆195Updated this week
- llm-d helm charts and deployment examples☆59May 1, 2026Updated 4 months ago
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆18Updated this week
- ☆28Updated this week
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,439Updated this week
- Variant optimization autoscaler for distributed inference workloads☆57Updated this week
- ☆22Mar 11, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Auto-tuning for vllm. Getting the best performance out of your LLM deployment (vllm+guidellm+optuna)☆64Aug 19, 2026Updated 2 weeks ago
- Gateway API Inference Extension☆760Updated this week
- Inference payload processor for llm-d☆19Updated this week
- ☆110Jul 21, 2025Updated last year
- Let my Claude talk to yours.☆30Aug 5, 2026Updated last month
- Inference Platform Simulation☆29Updated this week
- label ALL kubectl, kustomize, and helm objects, inline, without extra steps.(including namespaces and CRDs)☆15Apr 22, 2024Updated 2 years ago
- Queuing and quota management for AI/ML batch jobs on Kubernetes☆17Jul 9, 2026Updated last month
- Performance dashboards from the Perf & Scale team☆20Aug 15, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A tool to detect infrastructure issues on cloud native AI systems☆56Sep 18, 2025Updated 11 months ago
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Example app for the MCPI protocol☆10Apr 1, 2025Updated last year
- CLI for the Serverless Supercomputer☆25Sep 17, 2025Updated 11 months ago
- Offline optimization of your disaggregated Dynamo graph☆437Updated this week
- Discover ingress-nginx usage and auto-generate Gateway API migration plans before ingress-nginx reaches end-of-life (March 2026).☆16Nov 26, 2025Updated 9 months ago
- Community maintained hardware plugin for vLLM on Spyre☆53Aug 31, 2026Updated last week
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- ☆15May 28, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- kubeadm with core components instrumented to export OpenTelemetry traces (etcd, kube-apiserver, crio)☆13Oct 18, 2022Updated 3 years ago
- An ansible role which configures metrics collection.☆17Updated this week
- General agent evaluation framework☆76Aug 11, 2026Updated 3 weeks ago
- A workload for deploying LLM inference services on Kubernetes☆292Updated this week
- d.run website☆18Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- A framework for designing, executing and analysing experiment campaigns☆62Updated this week