LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
☆431Aug 6, 2026Updated this week
Alternatives and similar repositories for Lvllm
Users that are interested in Lvllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆105Updated this week
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- The complete NUMA-optimized branch of the ktransformers project☆25Nov 3, 2025Updated 9 months ago
- forked from vllm-project/flash-attention☆59May 9, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- fastllm是后端无依赖 的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆4,911Updated this week
- A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations☆19,215Updated this week
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode…☆483Jul 13, 2026Updated 3 weeks ago
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆586Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Aug 4, 2026Updated last week
- ☆43May 6, 2025Updated last year
- ☆16Updated this week
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆91Aug 3, 2026Updated last week
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆32Jul 2, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An extension utility for llama.cpp, used with 3090*2 + Strix Halo. llama.cpp的拓展小工具,自用于3090*2 + Strix Halo。☆310Updated this week
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- Run DeepSeek-V4-Flash on SM89 (Ada / RTX 4090) with vLLM — patch over PR #41834. Validated on 4x RTX 4090.☆95Updated this week
- 此仓库用于存放 eide 的相关二进制资源(注意:该仓库仅作为 eide 的下载站点,仅储存 zip, 7z 压缩包,且会定时清理 commit 记录,不必 clone 该仓库)☆11May 1, 2026Updated 3 months ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆31,578Updated this week
- The LLM API Benchmark Tool is a flexible Go-based utility designed to measure and analyze the performance of OpenAI-compatible API endpoi…☆96Mar 25, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆301Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆809Updated this week
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.☆7,998Updated this week
- FastAPI to serve Qwen-ASR with streaming support. Tested. Benchmarked. Flash Attention 2. Fast & Stable.☆15Jun 24, 2026Updated last month
- run DeepSeek-R1 GGUFs on KTransformers☆258Mar 3, 2025Updated last year
- aha model inference library, now supports Qwen(2.5VL/3/3VL/3.5/ASR/3Embedding/3Reranker), MiniCPM(4/5), VoxCPM(0.5B/1.5/2), DeepSeek-OCR/…☆383Jun 7, 2026Updated 2 months ago
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆61Updated this week
- ☆164Updated this week
- ☆16Jul 18, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆17Mar 20, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Aug 1, 2026Updated last week
- A high-performance Remodex WebSocket relay rewritten in Rust with Docker support☆46May 14, 2026Updated 2 months ago
- ☆79Apr 1, 2026Updated 4 months ago
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆182Jun 30, 2026Updated last month
- A high-throughput and memory-efficient inference and serving engine for LLMs☆88,592Updated this week
- A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.☆3,219Updated this week