Service-aware KV-cache compression for bandwidth-efficient disaggregated LLM serving.
☆16Jul 18, 2026Updated this week
Alternatives and similar repositories for KVServe
Users that are interested in KVServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Data on GPUs☆15Jun 21, 2026Updated 3 weeks ago
- Online Anomaly Detection for HPC Performance Data☆11Jun 25, 2018Updated 8 years ago
- DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression☆12Oct 7, 2020Updated 5 years ago
- COCCL: Compression and precision co-aware collective communication library☆36Jul 7, 2026Updated last week
- A portable implementation of SZ lossy compression for AMD GPUs and Hygon DCUs.☆11Feb 26, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- HDF5 Cache VOL connector for caching data on fast storage layers and moving data asynchronously to the parallel file system to hide I/O o…☆22Feb 10, 2026Updated 5 months ago
- ☆90Oct 17, 2025Updated 9 months ago
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆19Updated this week
- ☆20May 11, 2026Updated 2 months ago
- An artificial matrix generator in C☆13Feb 16, 2023Updated 3 years ago
- EGU26 SC2.5 Materials☆15Jun 29, 2026Updated 3 weeks ago
- ☆15Apr 11, 2024Updated 2 years ago
- ☆17Jun 12, 2026Updated last month
- Pressio is latin for compression. Libpressio is a C++ library with C compatible bindings to abstract between different lossless and lossy…☆16Dec 30, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆135Nov 11, 2024Updated last year
- An LLM inference engine, written in C++☆20Mar 30, 2026Updated 3 months ago
- Google DeepMind: Mixture of Depths Unofficial Implementation.☆12May 29, 2024Updated 2 years ago
- Some traditional multi-user communication methods as the baseline compared with semantic communication, including JPEG-LDPC-BPSK-NOMA, JP…☆21May 8, 2025Updated last year
- DLL注入工具☆13Nov 9, 2020Updated 5 years ago
- ☆25Oct 9, 2023Updated 2 years ago
- Artifacts accompanying the NSDI '24 paper: Leo: Online ML-based Traffic Classification at Multi-Terabit Line Rate.☆22Jun 13, 2026Updated last month
- Adaptive Entropy Coding library☆21Jan 27, 2022Updated 4 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5☆16Sep 19, 2024Updated last year
- Twenty Years After: Hierarchical Core-Stateless Fair Queueing☆17Feb 26, 2021Updated 5 years ago
- Unreal Engine 5 3D Platformer game prototype☆20May 27, 2024Updated 2 years ago
- A GPU accelerated error-bounded lossy compression for scientific data.☆100Jul 2, 2026Updated 2 weeks ago
- code repo for GCR [FAST'26]☆16Mar 3, 2026Updated 4 months ago
- PilotFish harvests the free GPU cycles of cloud gaming with deep learning training☆14Jul 2, 2022Updated 4 years ago
- Blazing Fast Erasure-Coding with Random Linear Network Coding (RLNC)☆46Oct 15, 2025Updated 9 months ago
- ☆50Dec 19, 2025Updated 7 months ago
- Official Implementation for (ICLR 2024) Idempotence and Perceptual Image Compression☆41May 29, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆47May 13, 2025Updated last year
- Implementation of TSM2L and TSM2R -- High-Performance Tall-and-Skinny Matrix-Matrix Multiplication Algorithms for CUDA☆35Jul 28, 2020Updated 5 years ago
- ☆32Apr 11, 2022Updated 4 years ago
- ☆40Nov 28, 2024Updated last year
- The official implementation of OSDI'25 paper BlitzScale☆48Apr 15, 2026Updated 3 months ago
- ☆18Sep 21, 2025Updated 9 months ago
- Parallel framework for training and fine-tuning deep neural networks☆74Apr 28, 2026Updated 2 months ago