A high-throughput and memory-efficient inference and serving engine for LLMs
☆49Sep 18, 2025Updated 10 months ago
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- verl: Volcano Engine Reinforcement Learning for LLMs☆42Jun 23, 2025Updated last year
- ☆16Nov 11, 2025Updated 9 months ago
- this is for fun, ain't it grand!☆22Sep 18, 2025Updated 10 months ago
- The official implementation for "ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning"☆15Jul 31, 2026Updated 2 weeks ago
- 使用fastrtc框架调用qwen-2.5-omni-realtime实现实时语音、视频等☆14Jun 27, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Introduction to Gaussian Processes☆29Jul 6, 2018Updated 8 years ago
- Load and visualize different datasets in video question answering☆10May 11, 2021Updated 5 years ago
- A curated list of research in machine learning system. I also summarize some papers if I think they are really interesting.☆11Nov 6, 2021Updated 4 years ago
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- ☆20Jul 11, 2025Updated last year
- 基于通义千问 Qwen2.5-Omni 的实时语音对话系统,使用在线API服务,支持实时语音交互、动态语音活动检测和流式音频处理。A real-time voice conversation system based on Qwen2.5-Omni Online-API, …☆92May 11, 2025Updated last year
- Knowledge-based robot flexible task planning project using classical planning, behavior trees and LLM techniques.☆84Dec 5, 2024Updated last year
- ☆23Mar 20, 2024Updated 2 years ago
- Material ListView and GridView using recycler view and card view☆10Sep 3, 2015Updated 10 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Nodes for ComfyUI to simply workflows☆76Jun 29, 2026Updated last month
- Invariant Feature Regularization for Fair Face Recognition (ICCV'23)☆15Oct 23, 2023Updated 2 years ago
- Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models (ACL-Findings 2024)☆16Apr 23, 2024Updated 2 years ago
- ☆13Jan 17, 2024Updated 2 years ago
- ComfyUI QwenVL and Qwen wrapper☆145Nov 29, 2025Updated 8 months ago
- 使用django对情感分析功能进行封装,里面包含使用情感词典和Bert模型进行情感分类,最后可以使用tensorFlow serving将模型部署在docker中运行。☆13Sep 23, 2019Updated 6 years ago
- ☆16Jul 12, 2024Updated 2 years ago
- Making ComfyUI more comfortable!☆19Sep 13, 2025Updated 11 months ago
- 中文 Python 笔记☆12Jan 15, 2018Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- PyTorch implementation for "Probabilistic Circuits for Variational Inference in Discrete Graphical Models", NeurIPS 2020☆17Oct 11, 2021Updated 4 years ago
- ComfyUI Nodes that integrate GLSL shader support.☆20Aug 25, 2025Updated 11 months ago
- Automatic defect recognition in X-ray testing using computer vision☆13Dec 8, 2018Updated 7 years ago
- ☆18Dec 7, 2023Updated 2 years ago
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆21Feb 9, 2026Updated 6 months ago
- [CVPR 2023] Bridging Precision and Confidence: A Train-Time Loss for Calibrating Object Detection☆31Jun 21, 2023Updated 3 years ago
- ComfyUI Node wrapper of SunoAI API☆21Dec 17, 2024Updated last year
- ☆13Apr 9, 2025Updated last year
- [NeurIPS 2024] SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words☆57Jun 25, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official implementation for "Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition" (AAAI 2022 Oral).☆17Mar 11, 2025Updated last year
- Provides path to various system directories.☆13Apr 25, 2025Updated last year
- Revision of official yolov7-pose to support custom dataset for keypoint detection☆11Nov 12, 2023Updated 2 years ago
- Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and pe…☆4,065Jun 12, 2025Updated last year
- Effective method to improve memory managment on Raspberry Pi and low RAM computers☆16Apr 17, 2020Updated 6 years ago
- This project is an SDK example for Azero developers to learn how to call the Linux version of Azero sdk API. You can run it quickly by fo…☆10Jul 31, 2020Updated 6 years ago
- ☆15Jul 23, 2025Updated last year