Predict the performance of LLM inference services
☆23Sep 18, 2025Updated last year
Alternatives and similar repositories for LLM-performance-prediction
Users that are interested in LLM-performance-prediction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- SpotServe: Serving Generative Large Language Models on Preemptible Instances☆138Feb 22, 2024Updated 2 years ago
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆25Nov 21, 2024Updated last year
- Serverless Paper Reading and Discussion☆37Jan 9, 2023Updated 3 years ago
- ☆185Mar 12, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Failure dataset accompanying the paper "How Bad Can a Bug Get? An Empirical Analysis of Software Failures in the OpenStack Cloud Computi…☆10Jun 12, 2020Updated 6 years ago
- LangBench applications and scripts☆14Jun 7, 2023Updated 3 years ago
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving (OSDI 23)☆94Jul 14, 2023Updated 3 years ago
- Official Repo for "SplitQuant / LLM-PQ: Resource-Efficient LLM Offline Serving on Heterogeneous GPUs via Phase-Aware Model Partition and …☆39Aug 29, 2025Updated last year
- Source code for OSDI 2023 paper titled "Cilantro - Performance-Aware Resource Allocation for General Objectives via Online Feedback"☆41Jul 6, 2023Updated 3 years ago
- Accurate, large-scale, and extensible simulator for LLM inference Systems☆681Aug 24, 2026Updated 3 weeks ago
- ☆20May 10, 2025Updated last year
- ☆17Apr 7, 2025Updated last year
- LLM Serving Performance Evaluation Harness☆84Feb 25, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆13Jun 20, 2025Updated last year
- ☆49Jun 27, 2024Updated 2 years ago
- A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems☆293Jun 30, 2026Updated 2 months ago
- ☆32May 28, 2024Updated 2 years ago
- Group administration repository for Tech: IOPMP Task Group☆13Dec 19, 2024Updated last year
- ☆18Oct 31, 2022Updated 3 years ago
- ☆16Jan 14, 2025Updated last year
- ☆20Sep 25, 2023Updated 2 years ago
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving (HPCA '23)☆14Jun 20, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- TraceWeaver is a research prototype for transparently tracing requests through a microservice without application instrumentation.☆23Sep 2, 2024Updated 2 years ago
- ☆47Jul 4, 2024Updated 2 years ago
- Implementing Wake on LAN in MAAS 2.2+☆15Oct 2, 2017Updated 8 years ago
- Releasing the spot availability traces used in "Can't Be Late" paper.☆27Mar 31, 2024Updated 2 years ago
- ☆10Dec 10, 2024Updated last year
- Official repository for paper "KeyEE: Enhancing Low-resource Generative Event Extraction with Auxiliary Keyword Sub-Prompt"☆10Jun 5, 2024Updated 2 years ago
- A library developed by Volcano Engine for high-performance reading and writing of PyTorch model files.☆27Jan 2, 2025Updated last year
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- ☆31Mar 20, 2022Updated 4 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [OSDI'24] Serving LLM-based Applications Efficiently with Semantic Variable☆226Sep 21, 2024Updated last year
- ☆27Aug 31, 2023Updated 3 years ago
- Dynamic batching library for Deep Learning inference. Tutorials for LLM, GPT scenarios.☆106Aug 24, 2026Updated 3 weeks ago
- APEX+ is an LLM Serving Simulator☆51Jun 16, 2025Updated last year
- ☆24May 14, 2025Updated last year
- Secure and Scalable Federated Learning using Serverless Computing☆14Jan 31, 2024Updated 2 years ago
- Awesome-papers is a collection of awesome papers about cloud computing including resource management, serverless, microservice, observer…☆128Dec 23, 2024Updated last year