Auto-tuning for vllm. Getting the best performance out of your LLM deployment (vllm+guidellm+optuna)
☆64Sep 25, 2026Updated this week
Alternatives and similar repositories for auto-tuning-vllm
Users that are interested in auto-tuning-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- helm charts for deploying models with llm-d☆32Updated this week
- A Python-based tool, trained on the state-of-the-art Google Pegasus model, specializing in generating abstracts from given YouTube video …☆10Aug 6, 2023Updated 3 years ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,650Updated this week
- A performance testing and analysis automation framework☆16Updated this week
- ✨ GenAIOps Enablement Lab Instructions - https://rhoai-genaiops.github.io/lab-instructions☆16Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- AI quickstart that provides interactive dashboard to analyze AI Model Performance as well as Openshift metrics collected from Prometheus☆27Jun 9, 2026Updated 3 months ago
- Synthetic Data Generation Toolkit for LLMs☆163Updated this week
- Models as a Service☆75Oct 21, 2025Updated 11 months ago
- ☆30Updated this week
- Simplified model deployment on llm-d☆28Jul 2, 2025Updated last year
- AI Agents Workshop with Red Hat AI☆14Feb 26, 2025Updated last year
- ☆27Jun 6, 2025Updated last year
- ☆19Updated this week
- Python library for Evaluation☆16Aug 21, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Collection of demos for building Llama Stack based apps on OpenShift☆64Updated this week
- Time series forecasting and analytics, powered by machine learning☆11Mar 25, 2025Updated last year
- Benchmark and optimize LLM inference across frameworks with ease☆201Jul 14, 2026Updated 2 months ago
- ☆59Aug 1, 2025Updated last year
- Inference Platform Simulation☆36Updated this week
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,660Updated this week
- AI-on-OpenShift website source code☆109Aug 24, 2026Updated last month
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual h…☆204Updated this week
- AIOps for Distributed Environments☆17Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- redis module unit tests with python (deprecated) please see RLTest☆12Sep 8, 2019Updated 7 years ago
- A lab/workshop for Red Hat OpenShift Data Science using simple fraud detection as an example workload☆22Jul 18, 2024Updated 2 years ago
- llm-d helm charts and deployment examples☆59May 1, 2026Updated 4 months ago
- An algorithm-focused interface for common llm training, continual learning, and reinforcement learning techniques☆98Updated this week
- Digital SuperTwin: digital twin of supercomputers☆13Nov 24, 2024Updated last year
- PM Workshop China☆10Apr 11, 2019Updated 7 years ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆854Updated this week
- Benchmark Suite Invocation Scripting☆11Mar 16, 2022Updated 4 years ago
- Agentic AI framework examples with Red Hat AI☆19Jul 2, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This project makes running the InstructLab large language model (LLM) fine-tuning process easy and flexible on OpenShift☆27Aug 27, 2025Updated last year
- Lola is able to package AI Context Modules or skills into a distributed package to be supported across multiple AI assistants. Think of y…☆126Updated this week
- ☆22Mar 11, 2026Updated 6 months ago
- ☆24Nov 18, 2025Updated 10 months ago
- Units of Measurement Libraries☆14Mar 2, 2026Updated 6 months ago
- A high performance batching router optimises max throughput for text inference workload☆16Sep 6, 2023Updated 3 years ago
- This is a clone of an SVN repository at http://pagecache-mangagement.googlecode.com/svn/trunk. It had been cloned by http://svn2github.co…☆10May 23, 2013Updated 13 years ago