A command-line interface tool for serving LLM using vLLM.
☆506Jan 25, 2026Updated 7 months ago
Alternatives and similar repositories for vllm-cli
Users that are interested in vllm-cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A CLI tool for managing your locally downloaded Huggingface models and datasets☆35Aug 19, 2025Updated last year
- A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes…☆521Apr 7, 2026Updated 4 months ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,557Updated this week
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 6 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,572Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆498Updated this week
- MMLU-Pro eval results☆15Aug 21, 2025Updated last year
- vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization☆2,538Updated this week
- ☆39Aug 4, 2025Updated last year
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,333Updated this week
- ☆46Feb 20, 2026Updated 6 months ago
- ☆2,793Jul 27, 2026Updated last month
- A high-performance and light-weight router for vLLM large scale deployment☆377Aug 18, 2026Updated last week
- Nano vLLM☆15,234Apr 26, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,743Updated this week
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,145Aug 23, 2026Updated last week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆32,926Updated this week
- ☆51Oct 1, 2025Updated 10 months ago
- Common recipes to run vLLM☆997Updated this week
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,682Updated this week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU cluste…☆5,185Aug 9, 2026Updated 3 weeks ago
- ☆24Dec 29, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A high-throughput and memory-efficient inference and serving engine for LLMs☆90,491Updated this week
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 10 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆783Updated this week
- vLLM Daily Summarization of Merged PRs☆53Updated this week
- MAESTRO is an AI-powered research application designed to streamline complex research tasks.☆1,494Apr 16, 2026Updated 4 months ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆467Updated this week
- ☆28May 10, 2026Updated 3 months ago
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.☆75,251Updated this week
- A framework for efficient model inference with omni-modality models☆6,478Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆483Nov 25, 2025Updated 9 months ago
- Renderer for the harmony response format to be used with gpt-oss☆4,493Apr 8, 2026Updated 4 months ago
- Recursive-Open-Meta-Agent v0.1 (Beta). A meta-agent framework to build high-performance multi-agent systems.☆5,176Feb 16, 2026Updated 6 months ago
- A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive vi…☆38,509Updated this week
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 4 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆7,021Aug 19, 2026Updated last week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,042Updated this week