Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
☆77Sep 11, 2026Updated this week
Alternatives and similar repositories for blackwell-llm-docker
Users that are interested in blackwell-llm-docker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆213Updated this week
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆92Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,038Updated this week
- ☆37Apr 26, 2026Updated 4 months ago
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- LuaJIT binding to Zydis - Zyantific disassembly tools☆16Aug 7, 2018Updated 8 years ago
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆26Sep 1, 2025Updated last year
- Example project using Zydis via git submodule and CMake☆17May 9, 2023Updated 3 years ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆541Updated this week
- These are performance benchmarks we did to prepare for our own privacy-preserving and NDA-compliant in-house AI coding assistant. If by a…☆32Apr 2, 2025Updated last year
- ☆12Sep 9, 2024Updated 2 years ago
- A proxy for minimax-m2, enabling interleaved thinking, and tool calls.☆39Nov 21, 2025Updated 9 months ago
- The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.☆14Mar 30, 2024Updated 2 years ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆164Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The sample project of Multi-Factor Authentication with ASP.NET Core 2.0 & Identity Server 4☆14Sep 21, 2018Updated 7 years ago
- Build different debian suites for the radxa rock 4 se☆12Mar 11, 2024Updated 2 years ago
- A ziglang implementation of the SSZ serialization protocol☆34Sep 1, 2026Updated last week
- ☆12Jan 29, 2024Updated 2 years ago
- Zydis Pascal Bindings☆22Nov 20, 2023Updated 2 years ago
- A lightweight, high-performance AI/RAG workspace and autonomous agent framework implemented in Go. Inspired by the Leann RAG backend arch…☆36Aug 17, 2026Updated 3 weeks ago
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 7 months ago
- Naming policies for the system JSON serializer in .NET.☆16Jun 18, 2025Updated last year
- Polkadot Standard Transactions Per Second (sTPS) performance benchmarking☆19Feb 3, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Zydis Python Bindings (Work In Progress)☆32Dec 20, 2021Updated 4 years ago
- ☆26May 2, 2026Updated 4 months ago
- An efficent mining turtle program for creating large underground builds in Minecraft.☆12Sep 3, 2025Updated last year
- 30 tok/s for 20B MoE on 8 GB VRAM. Flat throughput to 32K context. Native MXFP4 + GGUF Q4_K/Q5_K/Q6_K via ggml CUDA kernels — zero dequan…☆22Apr 7, 2026Updated 5 months ago
- Playable Bosses for Super Smash Bros. Ultimate.☆20Updated this week
- Fully local code indexing and sematic search tool☆16May 20, 2026Updated 3 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- ☆17Aug 13, 2026Updated last month
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 3 months ago
- Zydis Bindings for Go☆27Jul 18, 2021Updated 5 years ago
- Benchmarking tool for vLLM inference performance with GPU monitoring☆53Jun 7, 2026Updated 3 months ago
- Transport classes and utilities shared among .NET Elastic client libraries☆21Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- Serving and fine-tuning for the Gemma 4 and Qwen3.6 model families on Modal (SGLang/vLLM) — solo and concurrent shapes, custom chat-templ…☆32Jul 4, 2026Updated 2 months ago
- DGX Spark inference dashboard — vLLM, SGLang, llama.cpp, WebGPU & sparkrun with agent-powered Auto-Fix and speed optimization☆63Aug 3, 2026Updated last month