vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill
☆34Nov 7, 2025Updated 10 months ago
Alternatives and similar repositories for exvllm
Users that are interested in exvllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆462Sep 22, 2026Updated last week
- forked from vllm-project/flash-attention☆66May 9, 2026Updated 4 months ago
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆142Updated this week
- Run DeepSeek-V4.1-Flash / DeepSeek-V4-Flash and GLM-5.3-Flash on SM89 (Ada / RTX 4090) and SM120 (RTX PRO 6000) with vLLM☆177Updated this week
- DETR tensor去除推理过程无用辅助头+fp16部署再次加速+解决转tensorrt 输出全为0问题的新方法。☆12Jan 9, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Local LLM Inference Speed Test Tool☆231Sep 2, 2026Updated last month
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 3 months ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- yolo-pose for training escalator data☆16Jul 1, 2024Updated 2 years ago
- ☆974Sep 14, 2025Updated last year
- Towards Achieving Adversarial Robustness by Enforcing Feature Consistency Across Bit Planes☆23Jun 14, 2020Updated 6 years ago
- TensorRT-FastSAM(https://github.com/CASIA-IVA-Lab/FastSAM)☆23Feb 29, 2024Updated 2 years ago
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Vstream - Video Analytics pipeline with Hardware based accelerations (dev - stage)☆10Feb 2, 2024Updated 2 years ago
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,208Updated this week
- MobileSAM のエンコーダー/デコーダーをONNXに変換し、推論するサンプル☆12Apr 11, 2024Updated 2 years ago
- YOLOv12 TensorRT 端到端模型加速推理和INT8量化实现☆14Mar 5, 2025Updated last year
- Converts CLIP models to ONNX☆11Jan 17, 2023Updated 3 years ago
- A "standard library" of Triton kernels.☆26Oct 2, 2025Updated last year
- code☆13Jan 24, 2021Updated 5 years ago
- A C++ implementation for UCMCTrack (SOTA in MOT17)☆27May 30, 2025Updated last year
- Revision of official yolov7-pose to support custom dataset for keypoint detection☆11Nov 12, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An onnx-based quantitation tool.☆71Jan 8, 2024Updated 2 years ago
- LDC: Lightweight Dense CNN for Edge DetectionのPythonでのONNX推論サンプル☆15May 6, 2023Updated 3 years ago
- ☆10Feb 18, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,272Updated this week
- This Elgg plugin lets users preview MS Office files (doc, docx, xls, xlsx, ppt, pptx), Apple iWork pages, Adobe eps, and zip files using …☆12Aug 28, 2015Updated 11 years ago
- A simple panel to manage Linux☆23Jun 14, 2020Updated 6 years ago
- Deploy RT-EDTR with onnx from paddlepaddle framwork and graph cut☆32May 5, 2023Updated 3 years ago
- Multiple Lidar preprocessor for BEVfusion☆11Aug 25, 2023Updated 3 years ago
- 🎉My Collections of CUDA Kernels~☆11Jun 25, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆12Nov 26, 2024Updated last year
- Stable Diffusion in TensorRT 8.5+☆14Mar 19, 2023Updated 3 years ago
- 天池 NVIDIA TensorRT Hackathon 2023 —— 生成式AI模型优化赛 初赛第三名方案☆50Aug 16, 2023Updated 3 years ago
- An implementation of MSSRM method☆10Mar 23, 2023Updated 3 years ago
- Python scripts performing Open Vocabulary Object Detection using the YOLO-World model in ONNX. And Export the ONNX model for AXera's NPU☆12Aug 11, 2025Updated last year
- This repository provides tutorial, which discusses running sample publisher and subscriber using multiple transports of point_cloud_trans…☆11Aug 26, 2026Updated last month
- lightNet (Object Detection and Semantic Segmentation) for ONNX and TensorRT☆16Jul 4, 2023Updated 3 years ago