Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
☆142Sep 1, 2026Updated this week
Alternatives and similar repositories for modelexpress
Users that are interested in modelexpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆256Updated this week
- Benchmark SGLang on SLURM☆24Apr 20, 2026Updated 4 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆626Updated this week
- Offline optimization of your disaggregated Dynamo graph☆431Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,225Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆351Updated this week
- A workload for deploying LLM inference services on Kubernetes☆289Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- KV cache store for distributed LLM inference☆433Nov 13, 2025Updated 9 months ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆379Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,941Updated this week
- A toolkit for discovering cluster network topology.☆164Updated this week
- Speed up fsspec data access with Alluxio distributed caching.☆18Mar 22, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆160Aug 17, 2026Updated 2 weeks ago
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆18Updated this week
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆403Updated this week
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Aug 26, 2026Updated last week
- KJob: Tool for CLI-loving ML researchers☆44Jun 1, 2026Updated 3 months ago
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆502Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200…☆1,606Updated this week
- WG Serving☆38Mar 24, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Updated this week
- Kubernetes APIServer 高性能代理组件,代理 APIServer 的 List 请求,其它类型的请求会直接反向代理到原生 APIServer。 CKube 还额外支持了分页、搜索和索引等功能。 并且,CKube 100% 兼容原生 kubectl 和 ku…☆19Sep 16, 2022Updated 3 years ago
- helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/☆12Updated this week
- GenAI inference performance benchmarking tool☆237Updated this week
- A QA system based on k8s-specific knowledge build on ChatGLM2-6B, serving by Ray.☆10Sep 14, 2023Updated 2 years ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- LLM Kernel Library for Rust☆31Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 3 months ago
- High-performance safetensors model loader☆165Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSi…☆246Updated this week
- monitor InfiniBand usage by job or host☆27Jan 11, 2012Updated 14 years ago
- Gateway API Inference Extension☆758Updated this week
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 9 months ago
- Module, Model, and Tensor Serialization/Deserialization☆323Jul 7, 2026Updated last month
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆500Updated this week
- High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and S…☆200Updated this week