Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
☆86Jul 22, 2026Updated this week
Alternatives and similar repositories for modelexpress
Users that are interested in modelexpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆242Updated this week
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆461Updated this week
- Benchmark SGLang on SLURM☆24Apr 20, 2026Updated 3 months ago
- Offline optimization of your disaggregated Dynamo graph☆369Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,143Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆330Updated this week
- A workload for deploying LLM inference services on Kubernetes☆263Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- KV cache store for distributed LLM inference☆425Nov 13, 2025Updated 8 months ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆346Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,553Updated this week
- A toolkit for discovering cluster network topology.☆145Updated this week
- Speed up fsspec data access with Alluxio distributed caching.☆18Mar 22, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆151Updated this week
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆16Updated this week
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated 11 months ago
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆351Updated this week
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- KJob: Tool for CLI-loving ML researchers☆44Jun 1, 2026Updated last month
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆481Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B2…☆1,269Updated this week
- WG Serving☆38Mar 24, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Updated this week
- Kubernetes APIServer 高性能代理组件,代理 APIServer 的 List 请求,其它类型的请求会直接反向代理到原生 APIServer。 CKube 还额外支持了分页、搜索和索引等功能。 并且,CKube 100% 兼容原生 kubectl 和 ku…☆19Sep 16, 2022Updated 3 years ago
- helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/☆12Updated this week
- GenAI inference performance benchmarking tool☆212Updated this week
- A QA system based on k8s-specific knowledge build on ChatGLM2-6B, serving by Ray.☆10Sep 14, 2023Updated 2 years ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- LLM Kernel Library for Rust☆30Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 2 months ago
- High-performance safetensors model loader☆156Jul 7, 2026Updated 2 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSi…☆215Updated this week
- monitor InfiniBand usage by job or host☆26Jan 11, 2012Updated 14 years ago
- Gateway API Inference Extension☆723Updated this week
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 8 months ago
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆412Updated this week
- Module, Model, and Tensor Serialization/Deserialization☆318Jul 7, 2026Updated 2 weeks ago
- High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and S…☆183Updated this week