LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
☆457Sep 19, 2026Updated this week
Alternatives and similar repositories for Lvllm
Users that are interested in Lvllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆141Updated this week
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 10 months ago
- The complete NUMA-optimized branch of the ktransformers project☆26Nov 3, 2025Updated 10 months ago
- forked from vllm-project/flash-attention☆65May 9, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- fastllm是后端无依 赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆5,066Updated this week
- A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations☆19,524Updated this week
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-reques…☆906Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,102Updated this week
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 7 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,249Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- Unlock P2P comms between consumer NVIDIA GPUs☆539Updated this week
- ☆43May 6, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆17Aug 13, 2026Updated last month
- Qwen4-Exp (Qwen3.8-Flash-Next) SWA + MTP + PLE-on-iswa — bounded deep-context decode w/ long-range recall, TBQ4 4.25 bpv KV, DSpark + Rot…☆95Sep 2, 2026Updated 2 weeks ago
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆35Jul 2, 2026Updated 2 months ago
- An extension utility for llama.cpp, used with 3090*2 + Strix Halo. llama.cpp的拓展小工具,自用于3090*2 + Strix Halo。☆388Sep 7, 2026Updated 2 weeks ago
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆25Jan 30, 2026Updated 7 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,868Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆18Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,212Updated this week
- The LLM API Benchmark Tool is a flexible Go-based utility designed to measure and analyze the performance of OpenAI-compatible API endpoi…☆101Mar 25, 2026Updated 5 months ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆340Updated this week
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.☆8,083Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,080Updated this week
- FastAPI to serve Qwen-ASR with streaming support. Tested. Benchmarked. Flash Attention 2. Fast & Stable.☆16Jun 24, 2026Updated 2 months ago
- run DeepSeek-R1 GGUFs on KTransformers☆257Mar 3, 2025Updated last year
- aha model inference library, now supports Qwen(2.5VL/3/3VL/3.5/ASR/3Embedding/3Reranker), MiniCPM(4/5), VoxCPM(0.5B/1.5/2), DeepSeek-OCR/…☆394Jun 7, 2026Updated 3 months ago
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆82Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆254Updated this week
- ☆16Jul 18, 2023Updated 3 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Sep 11, 2026Updated last week
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆472Aug 17, 2026Updated last month
- ☆78Apr 1, 2026Updated 5 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆92,241Updated this week
- Sage attention for turning.☆82Dec 29, 2025Updated 8 months ago