Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troubleshooting
☆98Mar 11, 2026Updated 5 months ago
Alternatives and similar repositories for qwen3.5-dgx-spark
Users that are interested in qwen3.5-dgx-spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆123Feb 22, 2026Updated 6 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,235Updated this week
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆104Updated this week
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆106Oct 28, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆37Jan 6, 2026Updated 8 months ago
- ☆17Aug 13, 2026Updated 3 weeks ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆494Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆473Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated 3 months ago
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆458Jul 3, 2026Updated 2 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 7 months ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,332Updated this week
- Run vLLM on 1-to-N NVIDIA DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs☆129Jun 22, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, g…☆41Aug 11, 2026Updated 3 weeks ago
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆315Aug 23, 2026Updated 2 weeks ago
- ☆74Feb 27, 2026Updated 6 months ago
- comfyui optimizations for the dgx spark☆36Apr 30, 2026Updated 4 months ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆104Updated this week
- A MCP client that enables Mistral AI models to autonomously execute complex tasks across web and local environments through standardized …☆18Apr 1, 2025Updated last year
- Official Spark Arena Recipe Registry☆62Updated this week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official code for ICCV paper: MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh☆26Jul 9, 2026Updated last month
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆35Updated this week
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆33Mar 2, 2026Updated 6 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆111Updated this week
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆80Jul 18, 2026Updated last month
- A collection of templates/widgets for rapid prototyping☆12Jul 6, 2011Updated 15 years ago
- Crossplane Open Service Broker API☆21Updated this week
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆18Aug 10, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Feature flags for python.☆18May 22, 2017Updated 9 years ago
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆1,267Updated this week
- This is a FastAPI based LLM server. Load multiple LLM models (MLX or llama.cpp) simultaneously using multiprocessing.☆19Apr 8, 2026Updated 4 months ago
- The most advanced multisig bitcoin wallet optimized for collaboration.☆12Jun 11, 2026Updated 2 months ago
- Allows to parse CMakeLists.txt.☆13Apr 7, 2025Updated last year
- ☆15Nov 8, 2023Updated 2 years ago
- Code for "So similar and yet incompatible: Toward the automated identification of semantically compatible words" in NAACL 2015 proceedi…☆11May 11, 2015Updated 11 years ago