Unlock P2P comms between consumer NVIDIA GPUs
☆552Sep 16, 2026Updated last week
Alternatives and similar repositories for open-gpu-kernel-modules
Users that are interested in open-gpu-kernel-modules are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Practical local LLM recipes and benchmarks for RTX 5060 Ti setups☆171Sep 16, 2026Updated last week
- Make FP4 on 5090 Great Again☆24Jul 20, 2026Updated 2 months ago
- NVIDIA Linux open GPU with P2P support☆1,433Jun 6, 2025Updated last year
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,308Updated this week
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,530Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆446Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,260Updated this week
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆81Sep 19, 2026Updated last week
- ☆24Dec 29, 2025Updated 8 months ago
- Unified tracing profiler and visualizer, one timeline from CPU to GPU/HES to in-kernel zones☆147Updated this week
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆87Jul 12, 2026Updated 2 months ago
- ☆45Oct 9, 2025Updated 11 months ago
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆23Jun 20, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- easy exllama interface w/ automation & evals☆19Updated this week
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆101Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,168Updated this week
- ☆264Updated this week
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆142Sep 21, 2026Updated last week
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆461Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,883Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,109Updated this week
- The official API server for Exllama. OAI compatible, lightweight, and fast.☆1,434Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Core, junction, and VRAM temps for GDDR6, GDDR6X, GDDR7 NVIDIA GPUs in Linux☆97Updated this week
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆209Updated this week
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,126Updated this week
- An OpenAI API compatible FastAPI server that sits on top of the Anemll repo. Tested with Open WebUI.☆21Jan 21, 2026Updated 8 months ago
- Local LLM Inference Speed Test Tool☆230Sep 2, 2026Updated 3 weeks ago
- Tensor parallelism is all you need. Run LLMs on an AI cluster at home using any device. Distribute the workload, divide RAM usage, and in…☆18Nov 11, 2024Updated last year
- ☆32Jul 2, 2025Updated last year
- A harness optimized to smaller LLMs☆2,628Sep 18, 2026Updated last week
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆516Apr 17, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Qwen3-0.6B megakernel: 527 tok/s decode on RTX 3090 (3.8x faster than PyTorch)☆139Feb 10, 2026Updated 7 months ago
- ☆20Dec 24, 2024Updated last year
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.☆56Apr 5, 2026Updated 5 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- ☆59Oct 10, 2025Updated 11 months ago
- Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotunin…☆83Updated this week