Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
☆315Aug 23, 2026Updated 2 weeks ago
Alternatives and similar repositories for DGX_Spark_Qwen3.5-122B-A10B-AR-INT4
Users that are interested in DGX_Spark_Qwen3.5-122B-A10B-AR-INT4 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Docker configuration for running VLLM on dual DGX Sparks☆2,235Updated this week
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆60Jun 29, 2026Updated 2 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆104Updated this week
- Some benchmark results of small models and quants that fit on DGX Spark☆51Aug 23, 2026Updated 2 weeks ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆494Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆123Feb 22, 2026Updated 6 months ago
- Pure Rust Inference Engine☆676Updated this week
- Official Spark Arena Recipe Registry☆62Updated this week
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆391Aug 27, 2026Updated last week
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,332Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆473Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated 3 months ago
- ☆74Feb 27, 2026Updated 6 months ago
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆140Aug 26, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- llama-benchy - llama-bench style benchmarking tool for all backends☆694Jul 10, 2026Updated last month
- A model loader that uses fastsafetensors library to perform a fast, zero-copy load from storage to VRAM.☆23Jan 21, 2026Updated 7 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- ☆19May 31, 2026Updated 3 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- GLM-5.3 Flash EXL3 for 2x DGX Sparks☆374Updated this week
- ☆21Apr 7, 2026Updated 5 months ago
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆37Jan 6, 2026Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A monitor of resources for DGX Spark☆25Feb 13, 2026Updated 6 months ago
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆80Jun 28, 2026Updated 2 months ago
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆61Jul 30, 2026Updated last month
- An optimized setup for running ComfyUI on DGX Spark☆48Aug 2, 2026Updated last month
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆111Updated this week
- Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.☆80Jun 29, 2026Updated 2 months ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆25Jun 16, 2026Updated 2 months ago
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆98Mar 11, 2026Updated 5 months ago
- DGX Spark / GB10 vLLM image for Gemma 4 31B Deckard Heretic Uncensored NVFP4 with z-lab DFlash speculative decoding.☆33May 15, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vi…☆18Mar 12, 2026Updated 5 months ago
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆22Feb 16, 2026Updated 6 months ago
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆106Oct 28, 2025Updated 10 months ago
- DGX Spark inference dashboard — vLLM, SGLang, llama.cpp, WebGPU & sparkrun with agent-powered Auto-Fix and speed optimization☆61Aug 3, 2026Updated last month
- Lightweight nVidia telemetry and terminal system monitor - built for any architecture - Jetson, GB10, GB200, H100☆320Aug 18, 2026Updated 3 weeks ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 2 months ago
- ☆17Jun 27, 2026Updated 2 months ago