Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]
☆67Mar 5, 2025Updated last year
Alternatives and similar repositories for marconi
Users that are interested in marconi are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆21Jul 26, 2026Updated 3 weeks ago
- ☆17Aug 16, 2025Updated last year
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆75Oct 2, 2025Updated 10 months ago
- ☆36Jan 27, 2026Updated 6 months ago
- NEO is a LLM inference engine built to save the GPU memory crisis by CPU offloading☆99Jun 16, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A low-latency & high-throughput serving engine for LLMs☆518Jan 8, 2026Updated 7 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 8 months ago
- A Cluster-Wide Model Manager to Accelerate DNN Training via Automated Training Warmup☆36Jan 9, 2023Updated 3 years ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated 2 years ago
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆24Nov 21, 2024Updated last year
- Code and models for the paper: Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long …☆41Apr 9, 2026Updated 4 months ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- ☆21Jul 13, 2026Updated last month
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆119Dec 2, 2025Updated 8 months ago
- [ASPLOS'26] HILOS: A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs☆21Jan 18, 2026Updated 7 months ago
- continous batching and parallel acceleration for RWKV6☆23Jun 28, 2024Updated 2 years ago
- Expressive, Easy to Build, and High-Performance Application Networks☆20Jul 1, 2025Updated last year
- ☆39Nov 28, 2024Updated last year
- ☆17Dec 19, 2024Updated last year
- A throughput-oriented high-performance serving framework for LLMs☆975Mar 29, 2026Updated 4 months ago
- Efficient and easy multi-instance LLM serving☆564Mar 12, 2026Updated 5 months ago
- ☆41Nov 28, 2022Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 2 months ago
- ☆202Jul 15, 2025Updated last year
- [ICLR'25] "Understanding Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing" by Peihao Wang, Ruisi Cai, Yue…☆18Mar 21, 2025Updated last year
- ☆18Sep 21, 2025Updated 10 months ago
- ☆17May 10, 2024Updated 2 years ago
- The official implementation of OSDI'25 paper BlitzScale☆49Apr 15, 2026Updated 4 months ago
- A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of …☆332Jun 10, 2025Updated last year
- ATC23 AE☆45May 11, 2023Updated 3 years ago
- A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems☆286Jun 30, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Summary of some awesome work for optimizing LLM inference☆266Feb 14, 2026Updated 6 months ago
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆21Updated this week
- ☆11Oct 11, 2023Updated 2 years ago
- 🕹 Implementation for the lesson Compiling Engineering(2020 Spring) in Peking University, adjusted from UCLA CS 132 Project.☆10Jun 21, 2020Updated 6 years ago
- ☆20May 11, 2026Updated 3 months ago
- [NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning☆69Oct 31, 2025Updated 9 months ago
- ☆32May 28, 2024Updated 2 years ago