A high-throughput and memory-efficient inference and serving engine for LLMs
☆41Aug 13, 2026Updated last month
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Aug 3, 2026Updated last month
- ☆74Feb 27, 2026Updated 6 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- ☆17Aug 13, 2026Updated last month
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆19Aug 10, 2026Updated last month
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆29Sep 3, 2026Updated 2 weeks ago
- Open format for specifying structured assumptions and requirements about code.☆74Sep 11, 2026Updated last week
- ☆16Jun 3, 2026Updated 3 months ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,080Updated this week
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆26Aug 22, 2026Updated 3 weeks ago
- ☆42Jul 13, 2026Updated 2 months ago
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 华为2018 软件经验挑战赛 上合赛区决赛17名 227/300分☆12May 1, 2018Updated 8 years ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalit…☆33Jul 23, 2026Updated last month
- © 哨兵博客 V3 Power by Bin4xin | Jekyll | Github Action.☆11Updated this week
- ☆13Oct 27, 2019Updated 6 years ago
- Dockerfiles for poetry/mlc-llm(rk3588)/...☆10Sep 13, 2023Updated 3 years ago
- The continuous mountain car problem solved with DDPG☆13Apr 19, 2020Updated 6 years ago
- 一款被动扫描ssrf的burpsuite插件☆20Dec 30, 2022Updated 3 years ago
- The project “Behavioral Based Insider Threat Detection” leverages Deep learning to identify insider threats through user behavior and acc…☆11Sep 12, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 使用Sentencepiece对中文语料进行分词☆13Nov 30, 2023Updated 2 years ago
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆35Jul 2, 2026Updated 2 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,292Updated this week
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆36Dec 15, 2025Updated 9 months ago
- A compute framework for building Search, RAG, Recommendations and Analytics over complex (structured+unstructured) data, with ultra-modal…☆12Sep 16, 2024Updated 2 years ago
- ☆14Oct 22, 2023Updated 2 years ago
- Features and labels engineering of raw data of quotes of several stocks.☆32Oct 9, 2019Updated 6 years ago
- Kubeflow on OpenShift☆14Jan 24, 2019Updated 7 years ago
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Some benchmark results of small models and quants that fit on DGX Spark☆51Aug 23, 2026Updated 3 weeks ago
- 一款根据pom.xml获取引用的第三方组件的版本号并识别组件漏洞的工具☆23May 17, 2023Updated 3 years ago
- Hardware-adapted bridges that support devices using CAN protocol.☆21Jan 12, 2026Updated 8 months ago
- PydanticAI开源框架,搭建基于PostgreSQL、MySQL的Text2SQL应用进行SQL语句生成,支持GPT大模型、国产大模型、开源本地大模型☆18Dec 26, 2024Updated last year
- Indexed regex search for large codebases, powered by trigram / sparse n‑gram indexes. A grep-like CLI that builds a local on-disk index, …☆46Apr 30, 2026Updated 4 months ago
- ☆12Jun 17, 2023Updated 3 years ago
- ☆36May 20, 2024Updated 2 years ago