A high-throughput and memory-efficient inference and serving engine for LLMs
☆41Aug 13, 2026Updated 2 weeks ago
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Aug 3, 2026Updated 3 weeks ago
- ☆37Apr 26, 2026Updated 4 months ago
- ☆17Aug 13, 2026Updated 2 weeks ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆18Aug 6, 2026Updated 3 weeks ago
- ☆15May 24, 2021Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A framework and build automation tool to process exploits/payloads to evade antivirus and endpoint detection response products using reus…☆11Jan 16, 2024Updated 2 years ago
- Source code for MA4270: Data Modelling and Computation on Transformers and Nadaraya-Watson Kernel Regression☆19May 29, 2024Updated 2 years ago
- Simulates a logged in user.☆16Jul 10, 2024Updated 2 years ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆25Aug 22, 2026Updated last week
- ☆21Sep 27, 2023Updated 2 years ago
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆34Aug 17, 2026Updated 2 weeks ago
- GBM multicore scaling: h2o, xgboost and lightgbm on multicore and multi-socket systems☆20May 13, 2018Updated 8 years ago
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated 2 months ago
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalit…☆33Jul 23, 2026Updated last month
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- CVE-2021-4034 POC and Docker and Analysis write up☆12May 23, 2022Updated 4 years ago
- © 哨兵博客 V3 Power by Bin4xin | Jekyll | Github Action.☆11Updated this week
- An eclipse plugin for covert the encoding of files.☆19Mar 25, 2016Updated 10 years ago
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆34Jul 2, 2026Updated last month
- ☆194Updated this week
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆32Mar 2, 2026Updated 5 months ago
- Dockerfiles for poetry/mlc-llm(rk3588)/...☆10Sep 13, 2023Updated 2 years ago
- VDM sig bypass and additional WinAPI stubs☆20Feb 16, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆16May 7, 2025Updated last year
- 使用Sentencepiece对中文语料进行分词☆13Nov 30, 2023Updated 2 years ago
- Erebus is a payload generator written in Nim.☆18Jun 13, 2023Updated 3 years ago
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆35Dec 15, 2025Updated 8 months ago
- A compute framework for building Search, RAG, Recommendations and Analytics over complex (structured+unstructured) data, with ultra-modal…☆12Sep 16, 2024Updated last year
- ☆14Oct 22, 2023Updated 2 years ago
- Multi-model deliberation for pi, inspired by OpenRouter Fusion☆51Updated this week
- Features and labels engineering of raw data of quotes of several stocks.☆32Oct 9, 2019Updated 6 years ago
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Country flag FieldFormat Plugin for Kibana 7☆17Oct 23, 2020Updated 5 years ago
- Kubeflow on OpenShift☆14Jan 24, 2019Updated 7 years ago
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- 一款根据pom.xml获取引用的第三方组件的版本号并识别组件漏洞的工具☆23May 17, 2023Updated 3 years ago
- This repo consists of code for plotting top loss images☆13May 18, 2020Updated 6 years ago
- PydanticAI开源框架,搭建基于PostgreSQL、MySQL的Text2SQL应用进行SQL语句生成,支持GPT大模型、国产大模型、开源本地大模型☆18Dec 26, 2024Updated last year
- Docker configuration for running VLLM on dual DGX Sparks☆2,202Updated this week