☆22Mar 11, 2026Updated 4 months ago
Alternatives and similar repositories for inference-benchmark
Users that are interested in inference-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆40Updated this week
- ☆15May 11, 2025Updated last year
- ☆19Feb 18, 2026Updated 5 months ago
- ☆17Jan 23, 2026Updated 6 months ago
- ☆23Jul 7, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- torchprime is a reference model implementation for PyTorch on TPU.☆49Mar 3, 2026Updated 5 months ago
- GenAI inference performance benchmarking tool☆218Updated this week
- llm-d benchmark scripts and tooling☆64Updated this week
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 8 months ago
- Gateway API Inference Extension☆735Updated this week
- xpk (Accelerated Processing Kit, pronounced x-p-k,) is a software tool to help Cloud developers to orchestrate training jobs on accelerat…☆192Updated this week
- Aplus Framework Crypto Library☆17Mar 25, 2026Updated 4 months ago
- Recipes for reproducing training and serving benchmarks for large machine learning models using GPUs on Google Cloud.☆140Jul 28, 2026Updated last week
- ☆20Apr 16, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆121Updated this week
- A tool for coordinated checkpoint/restore of distributed applications with CRIU☆35Jul 31, 2026Updated last week
- ☆54Updated this week
- Package chanstream implements an API compatible with and similiar to the TCP connection (and net.Conn as well) API, on top of Go channels…☆14Sep 2, 2020Updated 5 years ago
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆17Updated this week
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆782Updated this week
- PyTorch/XLA integration with JetStream (https://github.com/google/JetStream) for LLM inference"☆86Dec 18, 2025Updated 7 months ago
- ☆29Apr 17, 2026Updated 3 months ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Package of Pathways-on-Cloud utilities☆32Updated this week
- Test Orchestrator for Performance and Scalability of AI pLatforms☆18Jun 23, 2026Updated last month
- An earning call robot built with LLM☆10Aug 4, 2023Updated 3 years ago
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- [Deprecated] Vulnerability scanner for containers and images☆13Oct 26, 2015Updated 10 years ago
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆487Updated this week
- Implementation of the LDP module block in PyTorch and Zeta from the paper: "MobileVLM: A Fast, Strong and Open Vision Language Assistant …☆15Mar 11, 2024Updated 2 years ago
- ☆11Apr 29, 2026Updated 3 months ago
- caniuse.com, but for kubernetes☆27Dec 25, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across speci…☆40Updated this week
- 中国开发者活动日程(关注点:开源、开发者、云原生)☆28Jul 31, 2026Updated last week
- JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs wel…☆451Jan 5, 2026Updated 7 months ago
- A toolkit for discovering cluster network topology.☆152Updated this week
- ☆17Apr 5, 2014Updated 12 years ago
- Helm charts for llm-d☆52Jul 22, 2025Updated last year
- A benchmarking tool for comparing different LLM API providers' DeepSeek model deployments.☆31Mar 28, 2025Updated last year