llama-server start/stop scripts for Qwen3.6-35B-A3B UD-Q8_K_XL GGUF on DGX Spark
☆27Jul 6, 2026Updated last month
Alternatives and similar repositories for Qwen3.6-35B-A3B-UD-Q8_K_XL_DGX-Spark-Recipe
Users that are interested in Qwen3.6-35B-A3B-UD-Q8_K_XL_DGX-Spark-Recipe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆60Jul 30, 2026Updated 3 weeks ago
- Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)☆34Jul 6, 2026Updated last month
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆280Updated this week
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆42Jul 13, 2026Updated last month
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆123Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆103Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆914Updated this week
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆80Jul 22, 2026Updated last month
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆155Updated this week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated last month
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆299Updated this week
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- macOS 26+ browser CLI for agents: persistent sessions, compact JSON, screenshots, clicks/forms, JS eval, and live preview in under 2 MB.☆67Jul 8, 2026Updated last month
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆422Updated this week
- Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.☆75Jun 29, 2026Updated last month
- BH hackathon☆14Apr 4, 2024Updated 2 years ago
- Let any AI agent run DaVinci Resolve — MCP server with live Studio control, free-edition FCPXML/LUT workflows, reference color matching, …☆31Jul 20, 2026Updated last month
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.uggested Donat…☆33Jul 20, 2026Updated last month
- Coding agent with cool features.☆19Jun 6, 2026Updated 2 months ago
- AI Context Takt: A pipeline tool for structural control over LLM context. Escape the black box of chat history, maximize token efficiency…☆21Jan 12, 2026Updated 7 months ago
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆33Aug 17, 2026Updated last week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- ☆17May 10, 2026Updated 3 months ago
- Screenplay breakdown & AI prompt-pack studio — import a script, get scenes, elements, bibles, shot lists & a timeline, then export self-e…☆42Aug 5, 2026Updated 2 weeks ago
- ☆15Jun 13, 2025Updated last year
- HermesHQ is a Docker-first control plane for running and operating multiple Hermes Agent instances from one web application.☆25Updated this week
- Local benchmarking UI for LLMs and AI agents☆23Apr 13, 2026Updated 4 months ago
- Performance-focused fork of Hermes Agent with side-by-side benchmarks, Turbo Score, perf dashboard, and token-savings reporting.☆20Aug 6, 2026Updated 2 weeks ago
- ☆13May 26, 2021Updated 5 years ago
- Headless 4K remote desktop for the NVIDIA DGX Spark (GB10): one-command installer for Sunshine + Moonlight low-latency game streaming wit…☆49Jun 3, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Local terminal-first coding agent with natural-language Work Sessions, bounded Forge execution, validation gates, and disposable worktree…☆31May 15, 2026Updated 3 months ago
- MLX binary vectors and associated algorithms.☆14Mar 13, 2025Updated last year
- Android MCP server - production-grade, token-optimized, accessibility-first☆16Jul 26, 2026Updated 3 weeks ago
- The end of Screenshot 2023-12-20-21.11.59.png☆15Dec 22, 2023Updated 2 years ago
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆80Jun 28, 2026Updated last month
- MMLU-Pro eval results☆15Aug 21, 2025Updated last year
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆71Jul 13, 2026Updated last month