Service-aware KV-cache compression for bandwidth-efficient disaggregated LLM serving.
☆22Jul 26, 2026Updated last month
Alternatives and similar repositories for KVServe
Users that are interested in KVServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Jun 12, 2026Updated 2 months ago
- Online Anomaly Detection for HPC Performance Data☆11Jun 25, 2018Updated 8 years ago
- DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression☆12Oct 7, 2020Updated 5 years ago
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- COCCL: Compression and precision co-aware collective communication library☆39Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A portable implementation of SZ lossy compression for AMD GPUs and Hygon DCUs.☆12Feb 26, 2025Updated last year
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆21Updated this week
- ☆21May 11, 2026Updated 3 months ago
- An artificial matrix generator in C☆13Feb 16, 2023Updated 3 years ago
- a library to characterize the data and check the compression results of lossy compressors☆20Aug 31, 2025Updated 11 months ago
- ☆15Apr 11, 2024Updated 2 years ago
- EGU26 SC2.5 Materials☆15Jun 29, 2026Updated 2 months ago
- An LLM inference engine, written in C++☆20Mar 30, 2026Updated 4 months ago
- Pressio is latin for compression. Libpressio is a C++ library with C compatible bindings to abstract between different lossless and lossy…☆16Dec 30, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆134Nov 11, 2024Updated last year
- Some traditional multi-user communication methods as the baseline compared with semantic communication, including JPEG-LDPC-BPSK-NOMA, JP…☆21Jul 28, 2026Updated last month
- DLL注入工具☆13Nov 9, 2020Updated 5 years ago
- ☆91Oct 17, 2025Updated 10 months ago
- Artifacts accompanying the NSDI '24 paper: Leo: Online ML-based Traffic Classification at Multi-Terabit Line Rate.☆22Jun 13, 2026Updated 2 months ago
- OpenGL 学习代码☆15Jun 25, 2023Updated 3 years ago
- Twenty Years After: Hierarchical Core-Stateless Fair Queueing☆18Feb 26, 2021Updated 5 years ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5☆16Sep 19, 2024Updated last year
- ☆15Jun 26, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Machine Learning Based DDoS Detection (HTTP,UDP,TCP and ICMP Flood Attack)☆10Jul 3, 2018Updated 8 years ago
- A GPU accelerated error-bounded lossy compression for scientific data.☆100Updated this week
- ☆24Updated this week
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems☆29Mar 3, 2026Updated 5 months ago
- Flow level simulation☆16Nov 22, 2015Updated 10 years ago
- LLM checkpointing for DeepSpeed/Megatron☆26Nov 30, 2025Updated 9 months ago
- ☆55Dec 19, 2025Updated 8 months ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆49May 13, 2025Updated last year
- 系统能力综合训练 riscv☆12Apr 24, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆39Nov 28, 2024Updated last year
- ☆18Sep 21, 2025Updated 11 months ago
- Parallel framework for training and fine-tuning deep neural networks☆75Apr 28, 2026Updated 4 months ago
- Implemented entropy-based detection using Python to allow POX controller to detect UDP Flood Attack in the simulated networks using Minin…☆22Dec 18, 2018Updated 7 years ago
- ☆34Feb 19, 2024Updated 2 years ago
- Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]☆67Mar 5, 2025Updated last year
- ☆168Apr 23, 2026Updated 4 months ago