A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively
☆125Feb 22, 2026Updated 7 months ago
Alternatives and similar repositories for dgx-vllm
Users that are interested in dgx-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Docker configuration for running VLLM on dual DGX Sparks☆2,339Updated this week
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆526Updated this week
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 6 months ago
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆315Aug 23, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆102Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆98Mar 11, 2026Updated 6 months ago
- Pure Rust Inference Engine☆700Updated this week
- ☆120Aug 6, 2026Updated last month
- Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for no…☆52Sep 20, 2026Updated last week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆469Sep 15, 2026Updated 2 weeks ago
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 3 months ago
- comfyui optimizations for the dgx spark☆36Apr 30, 2026Updated 5 months ago
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆59Jun 29, 2026Updated 3 months ago
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆354Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆74Feb 27, 2026Updated 7 months ago
- A model loader that uses fastsafetensors library to perform a fast, zero-copy load from storage to VRAM.☆23Jan 21, 2026Updated 8 months ago
- Run vLLM on 1-to-N NVIDIA DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs☆130Jun 22, 2026Updated 3 months ago
- llama-benchy - llama-bench style benchmarking tool for all backends☆740Jul 10, 2026Updated 2 months ago
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆128Sep 10, 2026Updated 3 weeks ago
- vLLM + Qwen3.5-122B-A10B-NVFP4 on NVIDIA DGX Spark (GB10/SM121) — single-GPU NVFP4 W4A4 with MTP speculative decoding, self-contained Doc…☆41Mar 12, 2026Updated 6 months ago
- An optimized setup for running ComfyUI on DGX Spark☆52Aug 2, 2026Updated 2 months ago
- ☆17Jan 15, 2026Updated 8 months ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,397Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- Setup guide for ML training on NVIDIA DGX Spark (GB10 Blackwell, CUDA 13, aarch64)☆182Feb 26, 2026Updated 7 months ago
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆33May 31, 2026Updated 4 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- It is an SFT Fine-tuning tool that performs no-code LLM fine-tuning for Nvidia DGX Spark and Asus Ascent GX10.☆45Apr 22, 2026Updated 5 months ago
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated last year
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆1,439Sep 26, 2026Updated last week
- Forage is for Storage☆12Jun 15, 2022Updated 4 years ago
- DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)☆28Aug 6, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Production-ready, bleeding-edge, upstream-first vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a), verified against real models…☆53Updated this week
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 10 months ago
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆24Aug 17, 2026Updated last month
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆37Jan 6, 2026Updated 8 months ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆27Aug 22, 2026Updated last month
- bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vi…☆18Mar 12, 2026Updated 6 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆81Jul 18, 2026Updated 2 months ago