Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troubleshooting
☆98Mar 11, 2026Updated 6 months ago
Alternatives and similar repositories for qwen3.5-dgx-spark
Users that are interested in qwen3.5-dgx-spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆125Feb 22, 2026Updated 7 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,324Updated this week
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆107Sep 5, 2026Updated 3 weeks ago
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆107Oct 28, 2025Updated 11 months ago
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆37Jan 6, 2026Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆522Updated this week
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆143Sep 12, 2026Updated 2 weeks ago
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆488Sep 10, 2026Updated 2 weeks ago
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆468Sep 15, 2026Updated last week
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 6 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 8 months ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,386Sep 10, 2026Updated 2 weeks ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆28Jun 16, 2026Updated 3 months ago
- Run vLLM on 1-to-N NVIDIA DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs☆130Jun 22, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆316Aug 23, 2026Updated last month
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆106Aug 20, 2026Updated last month
- Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, g…☆50Aug 11, 2026Updated last month
- ☆74Feb 27, 2026Updated 7 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 3 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- Official Spark Arena Recipe Registry☆63Sep 3, 2026Updated 3 weeks ago
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆145Jun 28, 2026Updated 3 months ago
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆50Aug 17, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆27Jul 23, 2026Updated 2 months ago
- Pure Rust Inference Engine☆700Updated this week
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆164Aug 29, 2026Updated 3 weeks ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆34Mar 2, 2026Updated 6 months ago
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆59Jun 29, 2026Updated 2 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆40Sep 21, 2026Updated last week
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆40Sep 20, 2026Updated last week
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆77Jul 25, 2026Updated 2 months ago
- Crossplane Open Service Broker API☆21Sep 15, 2026Updated last week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆19Aug 10, 2026Updated last month
- Stereo matching on two images for making a disparity map. My attempt to implement semi-global matching algorithm with some customs additi…☆10Mar 19, 2016Updated 10 years ago
- Feature flags for python.☆18May 22, 2017Updated 9 years ago
- Provides a `Project` CRD and controller for k8s to help with organising resources☆12Apr 19, 2024Updated 2 years ago
- Open format for specifying structured assumptions and requirements about code.☆92Sep 11, 2026Updated 2 weeks ago
- The most advanced multisig bitcoin wallet optimized for collaboration.☆12Jun 11, 2026Updated 3 months ago
- Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)☆35Jul 6, 2026Updated 2 months ago