DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)
☆37Jul 13, 2026Updated 2 months ago
Alternatives and similar repositories for DeepSeek-V4-Flash-2x-Spark-1M
Users that are interested in DeepSeek-V4-Flash-2x-Spark-1M are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆479Updated this week
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆106Aug 20, 2026Updated 3 weeks ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆82Jul 13, 2026Updated 2 months ago
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆77Jul 25, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.uggested Donat…☆34Jul 20, 2026Updated last month
- Coordinate local Codex sessions on macOS with project discovery, cooperative locks, and an optional Telegram bridge.☆24Updated this week
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆104Sep 5, 2026Updated last week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆26Jul 23, 2026Updated last month
- ☆56Updated this week
- CAD files for force feedback joystick base☆12Jun 19, 2021Updated 5 years ago
- my version of the models dev site which is way faster and has the filters I want☆18Apr 22, 2026Updated 4 months ago
- Open source, clear, transparent, real world llm benchmarks☆85Updated this week
- CAD files for force feedback joystick base☆17Oct 31, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 2 months ago
- ☆63Updated this week
- ☆39May 19, 2026Updated 3 months ago
- Native insanely fast terminal built for agents☆75Updated this week
- Video stabilization based on a user-defined feature with adaptive recognition using OpenCV☆20Aug 18, 2018Updated 8 years ago
- ☆88Updated this week
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated 3 weeks ago
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- A self-evolving agent in rust with one tool that only uses skills☆31Aug 12, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Engine monitoring for experimental aircraft on your tablet for about $150☆24May 24, 2018Updated 8 years ago
- ☆15Aug 6, 2021Updated 5 years ago
- A long-context code retrieval and reproduction benchmark.☆304Sep 4, 2026Updated last week
- This project provides a set of macros and sample config for Jubilee tool changers running on Klipper.☆19Oct 5, 2020Updated 5 years ago
- AI on consumer Blackwell — a field guide for the RTX 50-series (sm_120). Fixes, footguns, and honest measurement from one RTX 5090.☆25Updated this week
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆506Updated this week
- Turn GraphQL APIs into LLM Tools☆21Apr 7, 2025Updated last year
- An open-source hardware racing drone design and demonstration RL software for the artificial intelligence robotic drone race competitions☆13Sep 6, 2020Updated 6 years ago
- Claude Code plugin pattern: hook-intercepted commands that run code without an API call☆20Mar 31, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- 南京工业大学校园网自动登录脚本☆10Mar 21, 2023Updated 3 years ago
- AnsibleGo is a rewrite of ansible functionality in Golang for image building - with additional features and much less dependencies☆16Jan 29, 2024Updated 2 years ago
- Forgotten Developer - a terminal like theme for Frontity.☆11Jun 6, 2022Updated 4 years ago
- AI-powered text-to-video generation with dynamic prompt rewriting using Krea's realtime diffusion technology.☆17Dec 9, 2025Updated 9 months ago
- Learning Pytorch☆13Jun 12, 2018Updated 8 years ago
- Give any blind LLM eyes — on any machine — in one shot. Tiny local VLM (Qwen3.5-0.8B) as an OpenAI-compatible vision service. Mac (MLX) /…☆29Jul 7, 2026Updated 2 months ago
- Render *.drawio files to PNG or PDF files in a GitHub Action☆15Sep 18, 2023Updated 2 years ago