Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
☆2,228Sep 12, 2026Updated this week
Alternatives and similar repositories for club-3090
Users that are interested in club-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆119Apr 28, 2026Updated 4 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,851Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 4 months ago
- Single-file installer for a club-3090 webserver providing an admin control panel, a reverse proxy that automatically routes requests to t…☆25Jul 8, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,073Updated this week
- NVIDIA Linux open GPU with P2P support☆511Sep 4, 2026Updated last week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,650Updated this week
- vLLM Docker Container for Qwen3.6 27b☆51Jun 15, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,217Updated this week
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆271Updated this week
- A harness optimized to smaller LLMs☆2,579Aug 29, 2026Updated 2 weeks ago
- Experimental llama.cpp fork for inference research and development☆834Updated this week
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆227May 14, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Reproducible llama.cpp configs + per-category quality benches for Qwen3.6-27B on a single RTX 4090. Winners, dead ends, and the silent-co…☆24Apr 26, 2026Updated 4 months ago
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,374Updated this week
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆541Updated this week
- ☆20Jul 21, 2026Updated last month
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,038Updated this week
- ☆43May 4, 2026Updated 4 months ago
- ☆15Apr 27, 2026Updated 4 months ago
- AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.☆2,023Aug 12, 2026Updated last month
- LLM inference in C/C++☆2,374Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,076Aug 18, 2026Updated 3 weeks ago
- Structured Chain-of-Thought☆220May 16, 2026Updated 3 months ago
- Core, junction, and VRAM temps for GDDR6, GDDR6X, GDDR7 NVIDIA GPUs in Linux☆94Aug 29, 2026Updated 2 weeks ago
- An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, a…☆2,623Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆232Updated this week
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆74Apr 29, 2026Updated 4 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,764Sep 2, 2026Updated last week
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆80Jul 18, 2026Updated last month
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆84Jul 12, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆490Jun 22, 2026Updated 2 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆99Mar 19, 2026Updated 5 months ago
- Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore.☆3,216Updated this week
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆176Updated this week
- llama.cpp speculative decoding measured on one RTX 3090, Qwen3.6-35B-A3B UD-Q4_K_XL, commit 3737e4137. Published figures are re-derived f…☆35Sep 3, 2026Updated last week
- Using LLMs for iteratively exploring the solution search space at scale.☆740Sep 5, 2026Updated last week
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆521Sep 3, 2026Updated last week