LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
☆386Jul 21, 2026Updated this week
Alternatives and similar repositories for Lvllm
Users that are interested in Lvllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆97Updated this week
- 大语言模型工具集☆28Aug 1, 2025Updated 11 months ago
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 8 months ago
- The complete NUMA-optimized branch of the ktransformers project☆25Nov 3, 2025Updated 8 months ago
- forked from vllm-project/flash-attention☆58May 9, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Local LLM Inference Speed Test Tool☆116Jul 2, 2026Updated 2 weeks ago
- fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆4,855Updated this week
- A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations☆18,849Updated this week
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode…☆431Jul 13, 2026Updated last week
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆539Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆432Feb 20, 2026Updated 5 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆125Updated this week
- NVIDIA Linux open GPU with P2P support☆351Jul 13, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆43May 6, 2025Updated last year
- ☆15Updated this week
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆90May 17, 2026Updated 2 months ago
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆26Jul 2, 2026Updated 2 weeks ago
- KTransformers 一键部署脚本☆59Apr 18, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 5 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 7 months ago
- Flash Attention 2 implementation for Turing GPUs☆116Mar 23, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SGLang is a high-performance serving framework for large language models and multimodal models.☆30,583Updated this week
- The LLM API Benchmark Tool is a flexible Go-based utility designed to measure and analyze the performance of OpenAI-compatible API endpoi…☆90Mar 25, 2026Updated 3 months ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆286Jul 8, 2026Updated last week
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated last week
- Patched NVIDIA GRID Drivers☆36May 1, 2026Updated 2 months ago
- FastAPI to serve Qwen-ASR with streaming support. Tested. Benchmarked. Flash Attention 2. Fast & Stable.☆15Jun 24, 2026Updated 3 weeks ago
- run DeepSeek-R1 GGUFs on KTransformers☆258Mar 3, 2025Updated last year
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆41Updated this week
- ☆139Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆16Jul 18, 2023Updated 3 years ago
- ☆17Mar 20, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Jul 10, 2026Updated last week
- Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.☆293Updated this week
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆397May 29, 2026Updated last month
- ☆78Apr 1, 2026Updated 3 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 4 months ago