DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)
☆38Jul 13, 2026Updated last month
Alternatives and similar repositories for DeepSeek-V4-Flash-2x-Spark-1M
Users that are interested in DeepSeek-V4-Flash-2x-Spark-1M are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆43Jul 13, 2026Updated last month
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆21Aug 17, 2026Updated last week
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆124Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆440Updated this week
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆103Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆975Updated this week
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆81Jul 13, 2026Updated last month
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 655,360-token context + MTP spec decode (DCP4, fp8_ds_mla) on a 4x NVIDIA DGX Spark …☆30Jul 13, 2026Updated last month
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆75Jul 25, 2026Updated last month
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.uggested Donat…☆33Jul 20, 2026Updated last month
- Kubeflow on OpenShift☆14Jan 24, 2019Updated 7 years ago
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆102Jul 11, 2026Updated last month
- ☆41Updated this week
- ☆15Nov 10, 2020Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆80Jul 18, 2026Updated last month
- An example application connecting LoopBack and RethinkDB using loopback-connector-rethinkdb.☆12May 2, 2015Updated 11 years ago
- A plugin for Babel to redirect paths in import, export, require expressions☆11Jan 10, 2020Updated 6 years ago
- An (HTTP) API proxy server that stores and replays captured responses.☆12Feb 7, 2020Updated 6 years ago
- Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)☆35Jul 6, 2026Updated last month
- React library for invoking the context menu.☆12Mar 24, 2021Updated 5 years ago
- my version of the models dev site which is way faster and has the filters I want☆18Apr 22, 2026Updated 4 months ago
- Open source, clear, transparent, real world llm benchmarks☆72Updated this week
- Not a coding agent; a decoding agent☆36Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated last month
- Create REST resources and controllers with thinky and Express or Restify☆14Dec 14, 2016Updated 9 years ago
- Native insanely fast terminal built for agents☆75Updated this week
- Use borg and APFS snapshot to backup your Mac☆13Sep 18, 2019Updated 6 years ago
- ☆84Updated this week
- Helper utils and plugin for running unit tests on sketch plugins☆19Oct 28, 2017Updated 8 years ago
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated last week
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- A self-evolving agent in rust with one tool that only uses skills☆31Aug 12, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- My collection of dotfiles☆14Apr 22, 2026Updated 4 months ago
- Native MLX quants of Qwen3.6-27B-AEON-Ultimate-Uncensored for Apple Silicon (Metal): 8-bit, FP4, and MTP self-speculation. Preserves the …☆42Jun 28, 2026Updated last month
- ☆305Updated this week
- Identifies accessibility issues in your React.js elements☆19Oct 30, 2017Updated 8 years ago
- Remove the UUID in Node debugger URL☆16Mar 15, 2023Updated 3 years ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems☆468Updated this week
- ESP32 electronic leadscrew for a lathe☆28Oct 16, 2024Updated last year