A command-line interface tool for serving LLM using vLLM.
☆505Jan 25, 2026Updated 7 months ago
Alternatives and similar repositories for vllm-cli
Users that are interested in vllm-cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A CLI tool for managing your locally downloaded Huggingface models and datasets☆35Aug 19, 2025Updated last year
- A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes…☆532Apr 7, 2026Updated 5 months ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,630Updated this week
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 6 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,875Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆509Updated this week
- MMLU-Pro eval results☆15Aug 21, 2025Updated last year
- vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization☆2,621Updated this week
- ☆39Aug 4, 2025Updated last year
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,602Updated this week
- ☆46Feb 20, 2026Updated 7 months ago
- ☆2,798Jul 27, 2026Updated last month
- A high-performance and light-weight router for vLLM large scale deployment☆426Updated this week
- Nano vLLM☆15,535Apr 26, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,801Updated this week
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,445Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,212Updated this week
- Common recipes to run vLLM☆1,024Updated this week
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,761Updated this week
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 5 months ago
- The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU cluste…☆5,188Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆92,241Updated this week
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 10 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆847Updated this week
- vLLM Daily Summarization of Merged PRs☆54Updated this week
- MAESTRO is an AI-powered research application designed to streamline complex research tasks.☆1,492Apr 16, 2026Updated 5 months ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆485Updated this week
- ☆28May 10, 2026Updated 4 months ago
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.☆76,481Updated this week
- A framework for efficient model inference with omni-modality models☆6,932Updated this week
- ☆484Nov 25, 2025Updated 9 months ago
- Renderer for the harmony response format to be used with gpt-oss☆4,505Apr 8, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Recursive-Open-Meta-Agent v0.1 (Beta). A meta-agent framework to build high-performance multi-agent systems.☆5,180Feb 16, 2026Updated 7 months ago
- A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive vi…☆38,624Updated this week
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 4 months ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails…☆7,060Aug 19, 2026Updated last month
- Late Interaction Models Training & Retrieval☆894Jul 23, 2026Updated last month
- TokenSpeed is a speed-of-light LLM inference engine.☆2,151Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,460Updated this week