Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]
☆63Mar 5, 2025Updated last year
Alternatives and similar repositories for marconi
Users that are interested in marconi are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆20Updated this week
- ☆15Aug 16, 2025Updated 11 months ago
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆74Oct 2, 2025Updated 9 months ago
- ☆35Jan 27, 2026Updated 5 months ago
- NEO is a LLM inference engine built to save the GPU memory crisis by CPU offloading☆100Jun 16, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A low-latency & high-throughput serving engine for LLMs☆512Jan 8, 2026Updated 6 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆21Nov 28, 2025Updated 7 months ago
- A Cluster-Wide Model Manager to Accelerate DNN Training via Automated Training Warmup☆36Jan 9, 2023Updated 3 years ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated last year
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆24Nov 21, 2024Updated last year
- Code and models for the paper: Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long …☆40Apr 9, 2026Updated 3 months ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated 11 months ago
- ☆21Jul 13, 2026Updated last week
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated 11 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆115Dec 2, 2025Updated 7 months ago
- [ASPLOS'26] HILOS: A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs☆20Jan 18, 2026Updated 6 months ago
- continous batching and parallel acceleration for RWKV6☆23Jun 28, 2024Updated 2 years ago
- Expressive, Easy to Build, and High-Performance Application Networks☆20Jul 1, 2025Updated last year
- ☆40Nov 28, 2024Updated last year
- ☆17Dec 19, 2024Updated last year
- A throughput-oriented high-performance serving framework for LLMs☆970Mar 29, 2026Updated 3 months ago
- Efficient and easy multi-instance LLM serving☆564Mar 12, 2026Updated 4 months ago
- ☆41Nov 28, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆35May 22, 2026Updated 2 months ago
- ☆200Jul 15, 2025Updated last year
- [ICLR'25] "Understanding Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing" by Peihao Wang, Ruisi Cai, Yue…☆18Mar 21, 2025Updated last year
- ☆18Sep 21, 2025Updated 10 months ago
- ☆17May 10, 2024Updated 2 years ago
- The official implementation of OSDI'25 paper BlitzScale☆48Apr 15, 2026Updated 3 months ago
- A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of …☆330Jun 10, 2025Updated last year
- ATC23 AE☆45May 11, 2023Updated 3 years ago
- A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems☆282Jun 30, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Summary of some awesome work for optimizing LLM inference☆264Feb 14, 2026Updated 5 months ago
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆20Updated this week
- ☆11Oct 11, 2023Updated 2 years ago
- 🕹 Implementation for the lesson Compiling Engineering(2020 Spring) in Peking University, adjusted from UCLA CS 132 Project.☆10Jun 21, 2020Updated 6 years ago
- ☆20May 11, 2026Updated 2 months ago
- [NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning☆69Oct 31, 2025Updated 8 months ago
- ☆32May 28, 2024Updated 2 years ago