vLLM + Qwen3.5-122B-A10B-NVFP4 on NVIDIA DGX Spark (GB10/SM121) — single-GPU NVFP4 W4A4 with MTP speculative decoding, self-contained Docker build
☆41Mar 12, 2026Updated 6 months ago
Alternatives and similar repositories for SPARK_Qwen3.5-122B-A10B-NVFP4
Users that are interested in SPARK_Qwen3.5-122B-A10B-NVFP4 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM Qwen3.5-122B NVFP4 on DGX Spark (SM121) — full Docker build with 15 patches☆17Mar 17, 2026Updated 6 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Complete guide to running Qwen3.5-35B-A3B on NVIDIA DGX Spark (GB10) with vLLM - installation, benchmarks, vision features, and troublesh…☆98Mar 11, 2026Updated 6 months ago
- MiniMax M2 inference server for NVIDIA DGX Spark☆17Jan 24, 2026Updated 7 months ago
- A Chinese and English text translation plugin for ComfyUI.☆60Mar 31, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- a math-formula image recognition project which placed at the first place in a competition hosted by NAVER CONNECT boostcamp AI Tech☆10Dec 16, 2023Updated 2 years ago
- DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.☆58Sep 8, 2026Updated last week
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆125Feb 22, 2026Updated 6 months ago
- LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)☆21Dec 22, 2025Updated 8 months ago
- Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for no…☆52Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- Solver with Interface window for Cloudflare Turnstile and other Captchas.☆14Oct 7, 2024Updated last year
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆481Sep 10, 2026Updated last week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆27Jul 23, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 四足机器人控制板☆19Sep 26, 2022Updated 3 years ago
- RMON is a self-hosted uptime and synthetic monitoring platform for websites, APIs, servers, DNS and network services. Run checks from mul…☆25Updated this week
- Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, g…☆47Aug 11, 2026Updated last month
- Hopenet: deep head pose estimator on ncnn☆10Jun 18, 2020Updated 6 years ago
- Terminal voice-to-text TUI — Qwen3-ASR-1.7B on the Apple GPU via MLX (mlx-speech). Fully local, no PyTorch, transcribes in ~1s. macOS App…☆16Updated this week
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆315Aug 23, 2026Updated 3 weeks ago
- ☆39Jul 17, 2026Updated 2 months ago
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)☆106Oct 28, 2025Updated 10 months ago
- Open Source Auth Built on Freestyle: own your auth + data https://docs.freestyle.dev/guides/authentication/☆23Jun 12, 2024Updated 2 years ago
- Allow multiple clients to create task queues with different checkpoints☆20Jun 1, 2023Updated 3 years ago
- PiDiNet running in Android by ncnn☆15Sep 26, 2021Updated 4 years ago
- CLI tool that applies an ASCII filter to video or image.☆13Jun 20, 2023Updated 3 years ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆27Jun 16, 2026Updated 3 months ago
- Z.E.T.A. Zero: Cognitive Construct & Persistent Memory for Local LLMs☆46Feb 13, 2026Updated 7 months ago
- PoC code for CVE-2020-16939 Windows Group Policy DACL Overwrite Privilege Escalation☆12Oct 27, 2020Updated 5 years ago
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 2 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Add information from CDP or LLDP to SCCM Hardware Inventory☆16May 14, 2021Updated 5 years ago
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated 2 months ago
- Optimized pose detector inference for edge devices☆15Feb 23, 2023Updated 3 years ago
- DeepSeek-V4.1-Flash on 3-4x NVIDIA DGX Sparks☆138Updated this week
- A Keras-based recommendation engine for subreddits, channels on the popular social media site Reddit☆10Feb 24, 2024Updated 2 years ago
- ☆16Jan 30, 2020Updated 6 years ago
- free cohere☆13Aug 18, 2024Updated 2 years ago