Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSim), and more.
☆215Jul 20, 2026Updated this week
Alternatives and similar repositories for tair-kvcache
Users that are interested in tair-kvcache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆304Updated this week
- Offline optimization of your disaggregated Dynamo graph☆368Updated this week
- RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.☆1,282Updated this week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆5,925Updated this week
- KV cache store for distributed LLM inference☆425Nov 13, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- High Performance KV Cache Store for LLM☆59May 20, 2026Updated 2 months ago
- ☆150Apr 23, 2026Updated 2 months ago
- ☆377Jan 28, 2026Updated 5 months ago
- NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across speci…☆40Updated this week
- Agent skills for vLLM☆89Apr 3, 2026Updated 3 months ago
- mKernel: fast multi-node, multi-GPU fused kernels☆251Jun 21, 2026Updated last month
- Benchmark SGLang on SLURM☆24Apr 20, 2026Updated 3 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,139Updated this week
- ☆54Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆86Updated this week
- Accurate, large-scale, and extensible simulator for LLM inference Systems☆642Jul 25, 2025Updated 11 months ago
- ML kernels and benchmarking infrastructure written in TIRx☆66Updated this week
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated 11 months ago
- High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and S…☆179Updated this week
- ☆18Updated this week
- Persist and reuse KV Cache to speedup your LLM.☆302Updated this week
- LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure☆343Updated this week
- High Performance LLM Inference Operator Library☆1,041Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Muon in Int8 Precision Made Possible☆20Jun 18, 2026Updated last month
- Perplexity GPU Kernels☆591Nov 7, 2025Updated 8 months ago
- ☆159Updated this week
- ☆65Apr 26, 2025Updated last year
- Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B2…☆1,264Updated this week
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,467Updated this week
- Modular RDMA Interface☆151Updated this week
- Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serv…☆314Updated this week
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆100Apr 2, 2025Updated last year
- A workload for deploying LLM inference services on Kubernetes☆262Updated this week
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,573Jul 14, 2026Updated last week
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,106Updated this week
- A prefill & decode disaggregated LLM serving framework with shared GPU memory and fine-grained compute isolation.☆127Dec 25, 2025Updated 6 months ago
- A NCCL extension library, designed to efficiently offload GPU memory allocated by the NCCL communication library.☆110Dec 17, 2025Updated 7 months ago
- CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.☆534Updated this week