Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
☆162Sep 22, 2026Updated this week
Alternatives and similar repositories for modelexpress
Users that are interested in modelexpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆270Updated this week
- Benchmark SGLang on SLURM☆24Apr 20, 2026Updated 5 months ago
- Offline optimization of your disaggregated Dynamo graph☆446Updated this week
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆705Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,265Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆383Updated this week
- A workload for deploying LLM inference services on Kubernetes☆300Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- KV cache store for distributed LLM inference☆444Updated this week
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆387Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆8,145Updated this week
- A toolkit for discovering cluster network topology.☆171Updated this week
- Speed up fsspec data access with Alluxio distributed caching.☆18Mar 22, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆160Updated this week
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆23Updated this week
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆424Updated this week
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSi…☆260Updated this week
- KJob: Tool for CLI-loving ML researchers☆44Jun 1, 2026Updated 3 months ago
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆511Updated this week
- WG Serving☆38Mar 24, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200…☆1,750Updated this week
- Kubernetes APIServer 高性能代理组件,代理 APIServer 的 List 请求,其它类型的请求会直接反向代理到原生 APIServer。 CKube 还额外支持了分页、搜索和索引等功能。 并且,CKube 100% 兼容原生 kubectl 和 ku…☆19Sep 16, 2022Updated 4 years ago
- helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/☆13Updated this week
- GenAI inference performance benchmarking tool☆246Updated this week
- A QA system based on k8s-specific knowledge build on ChatGLM2-6B, serving by Ray.☆10Sep 14, 2023Updated 3 years ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆541Updated this week
- A high-performance and light-weight router for vLLM large scale deployment☆433Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LLM Kernel Library for Rust☆32Sep 9, 2026Updated last week
- llm-d Router: The intelligent entry point for inference requests☆353Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 4 months ago
- High-performance safetensors model loader☆167Sep 10, 2026Updated last week
- monitor InfiniBand usage by job or host☆27Jan 11, 2012Updated 14 years ago
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 10 months ago
- Gateway API Inference Extension☆770Updated this week