A command-line interface tool for serving LLM using vLLM.
☆507Jan 25, 2026Updated 6 months ago
Alternatives and similar repositories for vllm-cli
Users that are interested in vllm-cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A CLI tool for managing your locally downloaded Huggingface models and datasets☆35Aug 19, 2025Updated 11 months ago
- A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes…☆508Apr 7, 2026Updated 4 months ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,490Updated this week
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 5 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,093Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆487Updated this week
- MMLU-Pro eval results☆15Aug 21, 2025Updated 11 months ago
- vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization☆2,503Updated this week
- ☆39Aug 4, 2025Updated last year
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,004Updated this week
- ☆46Feb 20, 2026Updated 5 months ago
- ☆2,773Jul 27, 2026Updated 2 weeks ago
- A high-performance and light-weight router for vLLM large scale deployment☆351Updated this week
- Nano vLLM☆14,935Apr 26, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,652Updated this week
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,131Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆31,578Updated this week
- ☆51Oct 1, 2025Updated 10 months ago
- Common recipes to run vLLM☆963Updated this week
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,577Updated this week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU cluste…☆5,170Updated this week
- ☆24Dec 29, 2025Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A high-throughput and memory-efficient inference and serving engine for LLMs☆88,592Updated this week
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 9 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆711Updated this week
- vLLM Daily Summarization of Merged PRs☆53Updated this week
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆463Jul 14, 2026Updated 3 weeks ago
- MAESTRO is an AI-powered research application designed to streamline complex research tasks.☆1,494Apr 16, 2026Updated 3 months ago
- ☆28May 10, 2026Updated 3 months ago
- The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.☆69,824Updated this week
- A framework for efficient model inference with omni-modality models☆6,015Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆483Nov 25, 2025Updated 8 months ago
- Renderer for the harmony response format to be used with gpt-oss☆4,476Apr 8, 2026Updated 4 months ago
- Recursive-Open-Meta-Agent v0.1 (Beta). A meta-agent framework to build high-performance multi-agent systems.☆5,101Feb 16, 2026Updated 5 months ago
- A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive vi…☆38,013Jul 25, 2026Updated 2 weeks ago
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 3 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆6,992Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆1,840Updated this week