Predict the performance of LLM inference services
☆23Sep 18, 2025Updated 10 months ago
Alternatives and similar repositories for LLM-performance-prediction
Users that are interested in LLM-performance-prediction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SpotServe: Serving Generative Large Language Models on Preemptible Instances☆135Feb 22, 2024Updated 2 years ago
- Serverless Paper Reading and Discussion☆38Jan 9, 2023Updated 3 years ago
- Failure dataset accompanying the paper "How Bad Can a Bug Get? An Empirical Analysis of Software Failures in the OpenStack Cloud Computi…☆10Jun 12, 2020Updated 6 years ago
- LangBench applications and scripts☆14Jun 7, 2023Updated 3 years ago
- Simulator for the datacenter, including power, cooling, server and other components☆19Feb 12, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Releasing the spot availability traces used in "Can't Be Late" paper.☆26Mar 31, 2024Updated 2 years ago
- ☆20May 10, 2025Updated last year
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆24Nov 21, 2024Updated last year
- ☆13Jun 20, 2025Updated last year
- ☆18Oct 31, 2022Updated 3 years ago
- ☆179Mar 12, 2024Updated 2 years ago
- ☆20Sep 25, 2023Updated 2 years ago
- LLM Inference analyzer for different hardware platforms☆119Jun 23, 2026Updated 3 weeks ago
- Accurate, large-scale, and extensible simulator for LLM inference Systems☆642Jul 25, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆30Mar 20, 2022Updated 4 years ago
- ☆10Dec 10, 2024Updated last year
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- ☆21May 14, 2025Updated last year
- Dynamic batching library for Deep Learning inference. Tutorials for LLM, GPT scenarios.☆106Aug 14, 2024Updated last year
- A tool to detect infrastructure issues on cloud native AI systems☆53Sep 18, 2025Updated 10 months ago
- LLM Serving Performance Evaluation Harness☆84Feb 25, 2025Updated last year
- Official Tensorflow implementation for "Improving the Transferability of Adversarial Samples by Path-Augmented Method" (CVPR 2023).☆12Jun 16, 2023Updated 3 years ago
- Pytorch implementation for the pilot study on the robustness of latent diffusion models.☆13Jun 20, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This repository manifests set which is made to build a prototype system of TraceZip, made by 4 pieces.☆14Jul 17, 2025Updated last year
- CAShift: Benchmarking Log-Based Cloud Attack Detection under Normality Shift (FSE 2025)☆15Jun 25, 2026Updated 3 weeks ago
- E-commerce search benchmark is the first end-to-end application benchmark for e-commerce search system with personalized recommendations.…☆44Feb 15, 2023Updated 3 years ago
- A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems☆279Jun 30, 2026Updated 3 weeks ago
- Deduplication over dis-aggregated memory for Serverless Computing☆14Mar 21, 2022Updated 4 years ago
- ☆10Jun 4, 2024Updated 2 years ago
- Microsoft question-answering dataset☆10Jun 16, 2023Updated 3 years ago
- A Suite for Parallel Inference of Diffusion Transformers (DiTs) on multi-GPU Clusters☆58May 3, 2026Updated 2 months ago
- Serverless Apache Spark On AWS Fargate☆17Jun 1, 2019Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [OSDI'24] Serving LLM-based Applications Efficiently with Semantic Variable☆222Sep 21, 2024Updated last year
- ☆12Apr 23, 2026Updated 2 months ago
- ☆16Apr 13, 2024Updated 2 years ago
- [IJCAI'23] Speeding Up Multi-Objective Hyperparameter Optimization by Task Similarity-Based Meta-Learning for the Tree-Structured Parzen …☆10Apr 24, 2026Updated 2 months ago
- ☆13Jul 7, 2017Updated 9 years ago
- A platform that provides users with easy access to AI services developed by Montimage and usage of explainable AI techniques (e.g., LIME,…☆10Feb 17, 2026Updated 5 months ago
- [ICSE 2023] Differentiable interpretation and failure-inducing input generation for neural network numerical bugs.☆13Jan 5, 2024Updated 2 years ago