Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for now and single-spark. For the not-so-rich buddies. If you want latest/in-testing, look at the branches
☆51Sep 4, 2026Updated last week
Alternatives and similar repositories for dgx-spark-inference-stack
Users that are interested in dgx-spark-inference-stack are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A small Python library and CLI for working with Cloudflare Turnstile on your own pages - read the sitekey, build the token, verify the re…☆70Sep 6, 2026Updated last week
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆98Mar 11, 2026Updated 6 months ago
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆105Sep 5, 2026Updated last week
- NVIDIA DGX Spark Playbooks☆18Nov 26, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Headless 4K remote desktop for the NVIDIA DGX Spark (GB10): one-command installer for Sunshine + Moonlight low-latency game streaming wit…☆52Jun 3, 2026Updated 3 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,263Updated this week
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆115Updated this week
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆143Jun 28, 2026Updated 2 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 7 months ago
- GPU-accelerated WhisperX on NVIDIA Blackwell (SM_121) - DGX Spark compatible☆32Apr 23, 2026Updated 4 months ago
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆106Oct 28, 2025Updated 10 months ago
- Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, g…☆44Aug 11, 2026Updated last month
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆81Jun 28, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆23Feb 16, 2026Updated 6 months ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆505Updated this week
- Run Qwen3.5-35B-A3B with llama.cpp and openclaw on NVIDIA DGX Spark (GB10)☆72Mar 1, 2026Updated 6 months ago
- Some benchmark results of small models and quants that fit on DGX Spark☆51Aug 23, 2026Updated 2 weeks ago
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆398Aug 27, 2026Updated 2 weeks ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆80Jul 18, 2026Updated last month
- Manage your self-defined cheat sheets & knowledge base in Alfred☆14Dec 25, 2023Updated 2 years ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,353Updated this week
- Agentic CAD for 3D printing☆528Sep 5, 2026Updated last week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)☆21Dec 22, 2025Updated 8 months ago
- The world's smallest AI agent. ESP32 + pure C. $4 chip, ~120KB RAM.☆19Mar 4, 2026Updated 6 months ago
- Pragmatic Network Security for Cloud and Hybrid Networks☆10Nov 24, 2015Updated 10 years ago
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆314Aug 23, 2026Updated 3 weeks ago
- 📖 A repo of configuration examples for StackHawk's Hawkscan!☆18Jan 28, 2026Updated 7 months ago
- A curated collection of production-ready Hermes Agent skills — brainstorming, PRD workflows, debugging, Apple integrations, MLOps, docume…☆76May 2, 2026Updated 4 months ago
- Pure Rust Inference Engine☆679Updated this week
- A battle-tested solution for respawning node services on SysVinit-based (init.d) systems like Debian.☆15Mar 15, 2014Updated 12 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆13Jan 30, 2025Updated last year
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆31May 31, 2026Updated 3 months ago
- This repository contains examples of insecure code and/or security misconfigurations in common Agent frameworks☆20Apr 4, 2026Updated 5 months ago
- Setup guide for ML training on NVIDIA DGX Spark (GB10 Blackwell, CUDA 13, aarch64)☆181Feb 26, 2026Updated 6 months ago
- This repository contains custom Azure Policies for Azure Container Apps.☆12Sep 12, 2023Updated 3 years ago
- Chef Cookbook to install the Threat Stack agent☆17Aug 23, 2022Updated 4 years ago
- A model loader that uses fastsafetensors library to perform a fast, zero-copy load from storage to VRAM.☆23Jan 21, 2026Updated 7 months ago