A high-throughput and memory-efficient inference and serving engine for LLMs
☆37Aug 10, 2026Updated this week
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆35Apr 26, 2026Updated 3 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆102Updated this week
- The complete NUMA-optimized branch of the ktransformers project☆25Nov 3, 2025Updated 9 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆16May 11, 2026Updated 3 months ago
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆18Jul 27, 2026Updated 2 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Simulates a logged in user.☆16Jul 10, 2024Updated 2 years ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, ex…☆25Aug 3, 2026Updated last week
- PCB files for the Adafruit USB LiIon/LiPoly Charger☆18Jun 2, 2026Updated 2 months ago
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalit…☆32Jul 23, 2026Updated 2 weeks ago
- An Arduino shield for the MG2639 Cellular and GPS module.☆13Mar 29, 2018Updated 8 years ago
- 🔧 最新的基于Cesium 1.95与Three 143的整合示例☆11Jul 31, 2022Updated 4 years ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆30Mar 2, 2026Updated 5 months ago
- Dockerfiles for poetry/mlc-llm(rk3588)/...☆10Sep 13, 2023Updated 2 years ago
- VDM sig bypass and additional WinAPI stubs☆19Feb 16, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 使用Sentencepiece对中文语料进行分词☆13Nov 30, 2023Updated 2 years ago
- Stuff☆11Feb 19, 2014Updated 12 years ago
- 基于DC-SDK的开发的标绘工具,如点、线、面的绘制和一些军标的绘制☆13Feb 22, 2021Updated 5 years ago
- Erebus is a payload generator written in Nim.☆18Jun 13, 2023Updated 3 years ago
- AI文档校对工具是一个用于自动检查Word文档中错别字的应用程序。它通过基础标点规则将文档分割成短句,然后利用AI API进行错别字检查,并在原文档中添加批注指出错误。☆19May 6, 2025Updated last year
- ☆25Dec 21, 2016Updated 9 years ago
- Benchmark utils for Redis Cluster☆16Sep 20, 2018Updated 7 years ago
- Kubeflow on OpenShift☆14Jan 24, 2019Updated 7 years ago
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Experimentation on google's gemma model☆15Mar 6, 2024Updated 2 years ago
- PydanticAI开源框架,搭建基于PostgreSQL、MySQL的Text2SQL应用进行SQL语句生成,支持GPT大模型、国产大模型、开源本地大模型☆18Dec 26, 2024Updated last year
- ☆12Jun 17, 2023Updated 3 years ago
- Build a Full stack Q&A Chatbot with Langchain, and LLM Models on Amazon Sagemaker☆12Nov 10, 2023Updated 2 years ago
- It gives you a step by step approach to predict binary data using linear regression.☆11Feb 28, 2021Updated 5 years ago
- Deploy, launch and use LLMs on AWS☆16Jun 2, 2023Updated 3 years ago
- built using https://www.youtube.com/user/CS186Berkeley/playlists☆19Aug 14, 2020Updated 5 years ago
- Real-Time Voice AI Calling Assistant powered by Gemini Live API — native voice-to-voice, function calling, interrupt handling, zero backe…☆16Apr 11, 2026Updated 4 months ago
- Pip install yourself to a six figure career!☆12May 4, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Run a self-hosted Actions runner on Kubernetes.☆17Aug 20, 2020Updated 5 years ago
- REMERGE - Multi-Word Expression discovery algorithm☆15Jul 21, 2026Updated 3 weeks ago
- Portfolio of Allen Downey at Olin College☆21Dec 13, 2022Updated 3 years ago
- Data Science Take Home Challenges☆12Sep 21, 2018Updated 7 years ago
- Automated Machine Learning with Microsoft Azure, published by Packt☆23Apr 22, 2026Updated 3 months ago
- Tutorial for LLM developers about engine design, service deployment, evaluation/benchmark, etc. Provide a C/S style optimized LLM inferen…☆19Sep 5, 2023Updated 2 years ago
- ☆27Mar 15, 2023Updated 3 years ago