vLLM + Qwen3.5-122B-A10B-NVFP4 on NVIDIA DGX Spark (GB10/SM121) — single-GPU NVFP4 W4A4 with MTP speculative decoding, self-contained Docker build
☆41Mar 12, 2026Updated 6 months ago
Alternatives and similar repositories for SPARK_Qwen3.5-122B-A10B-NVFP4
Users that are interested in SPARK_Qwen3.5-122B-A10B-NVFP4 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM Qwen3.5-122B NVFP4 on DGX Spark (SM121) — full Docker build with 15 patches☆17Mar 17, 2026Updated 6 months ago
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆98Mar 11, 2026Updated 6 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 8 months ago
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆50Aug 17, 2026Updated last month
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆106Aug 20, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆126Feb 22, 2026Updated 7 months ago
- LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)☆22Dec 22, 2025Updated 9 months ago
- Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for no…☆52Updated this week
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆19Jul 13, 2026Updated 2 months ago
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆490Sep 10, 2026Updated last month
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆475Updated this week
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated last month
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆27Jul 23, 2026Updated 2 months ago
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆409Aug 27, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆17Aug 13, 2026Updated last month
- Run Qwen3.5-35B-A3B with llama.cpp and openclaw on NVIDIA DGX Spark (GB10)☆71Mar 1, 2026Updated 7 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 3 months ago
- Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, g…☆54Aug 11, 2026Updated last month
- Terminal voice-to-text TUI for Apple Silicon. Qwen3-ASR-1.7B or live-streaming Confucius4-R2T2 on the Apple GPU via MLX (mlx-speech). Ful…☆16Sep 27, 2026Updated last week
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆315Aug 23, 2026Updated last month
- ☆39Jul 17, 2026Updated 2 months ago
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆160Updated this week
- Pure Rust Inference Engine☆705Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆14Dec 1, 2025Updated 10 months ago
- Developed a Software for semantic segmentation of remote sensing imagery using Fully Convolutional Networks (FCNs). Initially, this softw…☆12Jul 19, 2021Updated 5 years ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 6 months ago
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆109Oct 28, 2025Updated 11 months ago
- Allow multiple clients to create task queues with different checkpoints☆20Jun 1, 2023Updated 3 years ago
- PiDiNet running in Android by ncnn☆15Sep 26, 2021Updated 5 years ago
- Bob 硅基流动 Siliconflow 语音合成插件☆16Apr 8, 2025Updated last year
- AIS decoder module for node.js☆12Feb 14, 2017Updated 9 years ago
- CLI tool that applies an ASCII filter to video or image.☆13Jun 20, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A modern shell with functional programming synatx.☆25Jul 11, 2026Updated 2 months ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆28Jun 16, 2026Updated 3 months ago
- PoC code for CVE-2020-16939 Windows Group Policy DACL Overwrite Privilege Escalation☆12Oct 27, 2020Updated 5 years ago
- Add information from CDP or LLDP to SCCM Hardware Inventory☆16May 14, 2021Updated 5 years ago
- NRL Simplified Multicast Forwarding (SMF) implementation (RFC 6621)☆13Sep 14, 2026Updated 3 weeks ago
- A docker container for FreeTakServer☆17Jul 9, 2020Updated 6 years ago
- Rotate Your Screen☆10Apr 27, 2019Updated 7 years ago