Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for now and single-spark. For the not-so-rich buddies. If you want latest/in-testing, look at the branches
☆51Jul 29, 2026Updated this week
Alternatives and similar repositories for dgx-spark-inference-stack
Users that are interested in dgx-spark-inference-stack are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆97Mar 11, 2026Updated 4 months ago
- Headless 4K remote desktop for the NVIDIA DGX Spark (GB10): one-command installer for Sunshine + Moonlight low-latency game streaming wit…☆44Jun 3, 2026Updated 2 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 6 months ago
- Single-file web UI for NVIDIA DGX Spark — pull Ollama models, browse and download from HuggingFace, manage LiteLLM routing, and control S…☆31May 19, 2026Updated 2 months ago
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆105Oct 28, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆19Feb 16, 2026Updated 5 months ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems☆421Updated this week
- Run Qwen3.5-35B-A3B with llama.cpp and openclaw on NVIDIA DGX Spark (GB10)☆72Mar 1, 2026Updated 5 months ago
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆219Updated this week
- Some benchmark results of small models and quants that fit on DGX Spark☆49Updated this week
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆79Jul 18, 2026Updated 2 weeks ago
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆54Jun 29, 2026Updated last month
- LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)☆21Dec 22, 2025Updated 7 months ago
- The world's smallest AI agent. ESP32 + pure C. $4 chip, ~120KB RAM.☆15Mar 4, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Pragmatic Network Security for Cloud and Hybrid Networks☆10Nov 24, 2015Updated 10 years ago
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 3 weeks ago
- 📖 A repo of configuration examples for StackHawk's Hawkscan!☆18Jan 28, 2026Updated 6 months ago
- Pure Rust Inference Engine☆627Updated this week
- Like vim-bootstrap, but for Kubernetes☆21Jun 1, 2026Updated 2 months ago
- A battle-tested solution for respawning node services on SysVinit-based (init.d) systems like Debian.☆15Mar 15, 2014Updated 12 years ago
- Dockerised Salt-Master☆16Jan 17, 2019Updated 7 years ago
- ☆13Jan 30, 2025Updated last year
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆24May 31, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM Stack for nVidia DGX Spark containing LiteLLM, LamaSwap, vLLM, Llama.cpp and ollama☆37Jul 6, 2026Updated 3 weeks ago
- mlx-lm server wrapper for agentic harness☆21Jan 26, 2026Updated 6 months ago
- This repository contains examples of insecure code and/or security misconfigurations in common Agent frameworks☆18Apr 4, 2026Updated 4 months ago
- Setup guide for ML training on NVIDIA DGX Spark (GB10 Blackwell, CUDA 13, aarch64)☆170Feb 26, 2026Updated 5 months ago
- Run vLLM on 1-to-N NVIDIA DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs☆124Jun 22, 2026Updated last month
- A model loader that uses fastsafetensors library to perform a fast, zero-copy load from storage to VRAM.☆21Jan 21, 2026Updated 6 months ago
- This GitHub repository contains a project that automates the provisioning of a Kubernetes (K8s) cluster using Infrastructure as Code (IaC…☆20Oct 19, 2025Updated 9 months ago
- Homebridge plugin for Pentair IntelliCenter☆12Apr 13, 2023Updated 3 years ago
- This GitHub repository contains all the code that is associated to the blog posts I have written.☆10Jan 15, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆16Updated this week
- Index your shell history in a full-text search database☆13Aug 23, 2024Updated last year
- Interface boards for the Neato LDS - the Lidar sensor of the XV-11.☆14Jun 21, 2014Updated 12 years ago
- ☆12Updated this week
- ☆14Apr 11, 2023Updated 3 years ago
- Outputs vs. outcomes: what's the different and why does it matter?☆17Apr 14, 2025Updated last year
- Security Scanning Samples with cnspec, cnquery, and Mondoo Platform☆18Jul 1, 2026Updated last month