Service-aware KV-cache compression for bandwidth-efficient disaggregated LLM serving.
☆46Sep 29, 2026Updated last week
Alternatives and similar repositories for KVServe
Users that are interested in KVServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Data on GPUs☆16Jun 21, 2026Updated 3 months ago
- ☆20Sep 19, 2026Updated 2 weeks ago
- Online Anomaly Detection for HPC Performance Data☆11Jun 25, 2018Updated 8 years ago
- DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression☆12Oct 7, 2020Updated 6 years ago
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- COCCL: Compression and precision co-aware collective communication library☆39Sep 21, 2026Updated 2 weeks ago
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆23Sep 29, 2026Updated last week
- ☆21May 11, 2026Updated 4 months ago
- An artificial matrix generator in C☆13Feb 16, 2023Updated 3 years ago
- RL-AFEC: Adaptive Forward Error Correction for Real-time Video Communication Based on Reinforcement Learning☆21Mar 31, 2022Updated 4 years ago
- a library to characterize the data and check the compression results of lossy compressors☆20Aug 31, 2025Updated last year
- An LLM inference engine, written in C++☆20Mar 30, 2026Updated 6 months ago
- Pressio is latin for compression. Libpressio is a C++ library with C compatible bindings to abstract between different lossless and lossy…☆16Dec 30, 2024Updated last year
- ☆137Nov 11, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Some traditional multi-user communication methods as the baseline compared with semantic communication, including JPEG-LDPC-BPSK-NOMA, JP…☆21Jul 28, 2026Updated 2 months ago
- DLL注入工具☆13Nov 9, 2020Updated 5 years ago
- ☆93Oct 17, 2025Updated 11 months ago
- Artifacts accompanying the NSDI '24 paper: Leo: Online ML-based Traffic Classification at Multi-Terabit Line Rate.☆22Jun 13, 2026Updated 3 months ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5☆16Sep 19, 2024Updated 2 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Machine Learning Based DDoS Detection (HTTP,UDP,TCP and ICMP Flood Attack)☆10Jul 3, 2018Updated 8 years ago
- Unreal Engine 5 3D Platformer game prototype☆22May 27, 2024Updated 2 years ago
- ☆25Oct 2, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems☆29Mar 3, 2026Updated 7 months ago
- code repo for GCR [FAST'26]☆17Mar 3, 2026Updated 7 months ago
- PilotFish harvests the free GPU cycles of cloud gaming with deep learning training☆14Jul 2, 2022Updated 4 years ago
- This is an official GitHub repository for the paper, "Towards timeout-less transport in commodity datacenter networks.".☆15Sep 7, 2022Updated 4 years ago
- pythonFlask+音视频SDK+OpenVino+OpenPose实现网课场景坐姿检测☆16Mar 20, 2023Updated 3 years ago
- Documentation for the Radar Chart Widget (Unreal Engine) Plugin☆20Jan 18, 2021Updated 5 years ago
- LLM checkpointing for DeepSpeed/Megatron☆27Nov 30, 2025Updated 10 months ago
- ☆21Jul 15, 2025Updated last year
- ☆57Dec 19, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Implementation for (ICLR 2024) Idempotence and Perceptual Image Compression☆41May 29, 2024Updated 2 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆49May 13, 2025Updated last year
- Implementation of TSM2L and TSM2R -- High-Performance Tall-and-Skinny Matrix-Matrix Multiplication Algorithms for CUDA☆35Jul 28, 2020Updated 6 years ago
- ☆38Nov 28, 2024Updated last year
- The official implementation of OSDI'25 paper BlitzScale☆50Apr 15, 2026Updated 5 months ago
- ☆18Sep 21, 2025Updated last year
- Parallel framework for training and fine-tuning deep neural networks☆76Apr 28, 2026Updated 5 months ago