Summary of some awesome work for optimizing LLM inference
☆264Feb 14, 2026Updated 5 months ago
Alternatives and similar repositories for LLM-inference-optimization-paper
Users that are interested in LLM-inference-optimization-paper are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆29Sep 26, 2025Updated 10 months ago
- Large Language Model (LLM) Systems Paper List☆2,220Jul 25, 2026Updated last week
- Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of pap…☆285Mar 6, 2025Updated last year
- Curated collection of papers in machine learning systems☆639Feb 7, 2026Updated 5 months ago
- 📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉☆5,446Jul 26, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This repository serves as a comprehensive survey of LLM development, featuring numerous research papers along with their corresponding co…☆344Jul 16, 2026Updated 3 weeks ago
- Curated collection of papers in MoE model inference☆414Mar 12, 2026Updated 4 months ago
- Disaggregated serving system for Large Language Models (LLMs).☆828Apr 6, 2025Updated last year
- Artifact of Chimera☆18May 6, 2025Updated last year
- Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]☆66Mar 5, 2025Updated last year
- ☆32May 28, 2024Updated 2 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure☆351Jul 29, 2026Updated last week
- ☆232Jul 27, 2026Updated last week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Accurate, large-scale, and extensible simulator for LLM inference Systems☆654Jul 25, 2025Updated last year
- A low-latency & high-throughput serving engine for LLMs☆514Jan 8, 2026Updated 6 months ago
- Efficient and easy multi-instance LLM serving