LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
☆451Aug 25, 2026Updated this week
Alternatives and similar repositories for Lvllm
Users that are interested in Lvllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆115Updated this week
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- forked from vllm-project/flash-attention☆64May 9, 2026Updated 3 months ago
- Local LLM Inference Speed Test Tool☆190Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆4,947Updated this week
- A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations☆19,329Updated this week
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-reques…☆742Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆718Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 6 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Aug 24, 2026Updated last week
- NVIDIA Linux open GPU with P2P support☆459Aug 22, 2026Updated last week
- ☆43May 6, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded …☆92Updated this week
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆34Jul 2, 2026Updated last month
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆16Jul 1, 2026Updated last month
- An extension utility for llama.cpp, used with 3090*2 + Strix Halo. llama.cpp的拓展小工具,自用于3090*2 + Strix Halo。☆350Updated this week
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- Run DeepSeek-V4-Flash GLM-5.3-Flash on SM89 (Ada / RTX 4090) SM120(RTX Pro 6000) with vLLM☆134Updated this week
- Flash Attention 2 implementation for Turing GPUs☆125Mar 23, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- SGLang is a high-performance serving framework for large language models and multimodal models.☆32,926Updated this week
- A prompt-aware LLM router that predicts which models can complete each request, then selects the cheapest capable one: 53.2% lower cost a…☆37Updated this week
- The LLM API Benchmark Tool is a flexible Go-based utility designed to measure and analyze the performance of OpenAI-compatible API endpoi…☆100Mar 25, 2026Updated 5 months ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆327Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆903Updated this week
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.☆8,035Updated this week
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated last month
- FastAPI to serve Qwen-ASR with streaming support. Tested. Benchmarked. Flash Attention 2. Fast & Stable.☆15Jun 24, 2026Updated 2 months ago
- run DeepSeek-R1 GGUFs on KTransformers☆257Mar 3, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- aha model inference library, now supports Qwen(2.5VL/3/3VL/3.5/ASR/3Embedding/3Reranker), MiniCPM(4/5), VoxCPM(0.5B/1.5/2), DeepSeek-OCR/…☆389Jun 7, 2026Updated 2 months ago
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆69Updated this week
- ☆194Updated this week
- ☆16Jul 18, 2023Updated 3 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week
- Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.☆309Updated this week
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆460Aug 17, 2026Updated 2 weeks ago