Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
☆113Aug 11, 2026Updated this week
Alternatives and similar repositories for modelexpress
Users that are interested in modelexpress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆251Updated this week
- Benchmark SGLang on SLURM☆24Apr 20, 2026Updated 3 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆532Updated this week
- Offline optimization of your disaggregated Dynamo graph☆403Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,185Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆336Updated this week
- A workload for deploying LLM inference services on Kubernetes☆274Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- KV cache store for distributed LLM inference☆425Nov 13, 2025Updated 8 months ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆367Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,738Updated this week
- A toolkit for discovering cluster network topology.☆154Updated this week
- Speed up fsspec data access with Alluxio distributed caching.☆18Mar 22, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆153Updated this week
- Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping☆17Updated this week
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆382Updated this week
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Jul 23, 2026Updated 2 weeks ago
- KJob: Tool for CLI-loving ML researchers☆44Jun 1, 2026Updated 2 months ago
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆490Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200…☆1,372Updated this week
- WG Serving☆38Mar 24, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆21Updated this week
- Kubernetes APIServer 高性能代理组件,代理 APIServer 的 List 请求,其它类型的请求会直接反向代理到原生 APIServer。 CKube 还额外支持了分页、搜索和索引等功能。 并且,CKube 100% 兼容原生 kubectl 和 ku…☆19Sep 16, 2022Updated 3 years ago
- helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/☆12Updated this week
- GenAI inference performance benchmarking tool☆218Updated this week
- A QA system based on k8s-specific knowledge build on ChatGLM2-6B, serving by Ray.☆10Sep 14, 2023Updated 2 years ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- LLM Kernel Library for Rust☆30Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 2 months ago
- High-performance safetensors model loader☆163Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSi…☆224Updated this week
- monitor InfiniBand usage by job or host☆26Jan 11, 2012Updated 14 years ago
- Gateway API Inference Extension☆738Updated this week
- Incubating P/D sidecar for llm-d☆17Nov 13, 2025Updated 8 months ago
- Module, Model, and Tensor Serialization/Deserialization☆320Jul 7, 2026Updated last month
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆457Updated this week
- High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and S…☆186Updated this week