☆22Mar 11, 2026Updated 6 months ago
Alternatives and similar repositories for inference-benchmark
Users that are interested in inference-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆42Oct 1, 2026Updated last week
- ☆15May 11, 2025Updated last year
- ☆17Jan 23, 2026Updated 8 months ago
- torchprime is a reference model implementation for PyTorch on TPU.☆49Mar 3, 2026Updated 7 months ago
- GenAI inference performance benchmarking tool☆256Oct 2, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Feb 18, 2025Updated last year
- llm-d benchmark scripts and tooling☆72Updated this week
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 10 months ago
- Gateway API Inference Extension☆782Updated this week
- xpk (Accelerated Processing Kit, pronounced x-p-k,) is a software tool to help Cloud developers to orchestrate training jobs on accelerat…☆194Aug 27, 2026Updated last month
- Recipes for reproducing training and serving benchmarks for large machine learning models using GPUs on Google Cloud.☆141Sep 8, 2026Updated last month
- ☆128Oct 2, 2026Updated last week
- A tool for coordinated checkpoint/restore of distributed applications with CRIU☆35Jul 31, 2026Updated 2 months ago
- ☆93Oct 2, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆850Updated this week
- Machine Learning from Human Preferences☆41Mar 23, 2026Updated 6 months ago
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆25Updated this week
- ☆13Oct 27, 2023Updated 2 years ago
- A starting point for creating service brokers implementing the Open Service Broker API☆30Aug 11, 2017Updated 9 years ago
- ☆18Feb 17, 2020Updated 6 years ago
- Compatibility layer to run Cloud Foundry applications on OpenShift☆11Sep 2, 2016Updated 10 years ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- Test Orchestrator for Performance and Scalability of AI pLatforms☆18Jun 23, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Collection of LLM completions for reasoning-gym task datasets☆31Jul 4, 2025Updated last year
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- [Deprecated] Vulnerability scanner for containers and images☆13Oct 26, 2015Updated 10 years ago
- A collection of useful Go libraries to ease the development of NVIDIA Operators for GPU/NIC management.☆30Updated this week
- Mixtral-based Ja-En (En-Ja) Translation model☆21Jan 6, 2025Updated last year
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆513Updated this week
- Synchronizes OpenShift BuildConfig objects as Jenkins jobs and synchronizes job status into OpenShift Build objects☆17Mar 24, 2026Updated 6 months ago
- Jax implementation of a flexible audio loudness meter in Python with implementation of ITU-R BS.1770-4 loudness algorithm☆13Jan 29, 2025Updated last year
- Implementation of the LDP module block in PyTorch and Zeta from the paper: "MobileVLM: A Fast, Strong and Open Vision Language Assistant …☆15Mar 11, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- caniuse.com, but for kubernetes☆27Dec 25, 2024Updated last year
- ☆34Nov 4, 2024Updated last year
- The original Shared Recurrent Memory Transformer implementation☆39Aug 24, 2026Updated last month
- 中国开发者活动日程(关注点:开源、开发者、云原生)☆30Sep 28, 2026Updated last week
- ☆93Updated this week
- JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs wel…☆464Jan 5, 2026Updated 9 months ago
- Command-line tools for managing OCI model artifacts, which are bundled based on Model Spec☆81Updated this week