Auto-tuning for vllm. Getting the best performance out of your LLM deployment (vllm+guidellm+optuna)
☆64Jun 12, 2026Updated 2 months ago
Alternatives and similar repositories for auto-tuning-vllm
Users that are interested in auto-tuning-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- helm charts for deploying models with llm-d☆32Jul 25, 2026Updated 3 weeks ago
- llm-d benchmark scripts and tooling☆64Updated this week
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,523Updated this week
- Starter kits for building and deploying AI agents. Run interactively locally or deploy to Red Hat OpenShift (including RHOAI) via OGX.☆27Updated this week
- A performance testing and analysis automation framework☆16Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Synthetic Data Generation Toolkit for LLMs☆157Updated this week
- The AI Accelerator is a template project for setting up Red Hat OpenShift AI using GitOps☆70Aug 11, 2026Updated last week
- ☆27Updated this week
- Simplified model deployment on llm-d☆28Jul 2, 2025Updated last year
- AI Agents Workshop with Red Hat AI☆13Feb 26, 2025Updated last year
- ☆19Jul 26, 2026Updated 3 weeks ago
- Python library for Evaluation☆17Mar 31, 2026Updated 4 months ago
- Vanilla configurations for a RHOAI instance to deploy a GenAI POC. This will deploy a vector database (Milvus), a GenAI interface (Anythi…☆14Apr 18, 2025Updated last year
- Inference Platform Simulation☆24Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Time series forecasting and analytics, powered by machine learning☆11Mar 25, 2025Updated last year
- Benchmark and optimize LLM inference across frameworks with ease☆200Jul 14, 2026Updated last month
- ☆59Aug 1, 2025Updated last year
- Operator for the OpenShift Lightspeed Service☆15Updated this week
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,049Updated this week
- AI-on-OpenShift website source code☆108Updated this week
- 100 days of LLM inference engineering — daily posts, experiments, and visualizations☆587Apr 30, 2026Updated 3 months ago
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual h…☆185Updated this week
- AIOps for Distributed Environments☆16Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- GenAI inference performance benchmarking tool☆222Updated this week
- redis module unit tests with python (deprecated) please see RLTest☆12Sep 8, 2019Updated 6 years ago
- A lab/workshop for Red Hat OpenShift Data Science using simple fraud detection as an example workload☆21Jul 18, 2024Updated 2 years ago
- llm-d helm charts and deployment examples☆59May 1, 2026Updated 3 months ago
- Red Hat Ecosystem Engineering - Agentic Collections☆49Aug 7, 2026Updated last week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆739Updated this week
- This project makes running the InstructLab large language model (LLM) fine-tuning process easy and flexible on OpenShift☆27Aug 27, 2025Updated 11 months ago
- SPDK fork of nvme-cli. No longer supported - use standard nvme-cli with SPDK nvme CUSE instead. See https://spdk.io/doc/nvme.html#nvme_…☆15Apr 10, 2024Updated 2 years ago
- Lola is able to package AI Context Modules or skills into a distributed package to be supported across multiple AI assistants. Think of y…☆117Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆24Nov 18, 2025Updated 9 months ago
- Units of Measurement Libraries☆14Mar 2, 2026Updated 5 months ago
- htop.dev website☆13Nov 5, 2024Updated last year
- A high performance batching router optimises max throughput for text inference workload☆16Sep 6, 2023Updated 2 years ago
- ODH Tools & Extensions Companion☆31Feb 8, 2026Updated 6 months ago
- This is a clone of an SVN repository at http://pagecache-mangagement.googlecode.com/svn/trunk. It had been cloned by http://svn2github.co…☆10May 23, 2013Updated 13 years ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆551Updated this week