MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 beats thinking-ON 90.6 for tool/agent work
☆39Jul 13, 2026Updated last month
Alternatives and similar repositories for MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark
Users that are interested in MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆42Jul 13, 2026Updated last month
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated last week
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆914Updated this week
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆103Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆20Aug 17, 2026Updated last week
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆60Jul 30, 2026Updated 3 weeks ago
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆80Jul 22, 2026Updated last month
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆123Updated this week
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆299Updated this week
- Run Faster-Qwen3-TTS on NVIDIA DGX Spark GB10 (ARM64/SM121/CUDA13) - OpenAI-compatible TTS API with CUDA graph acceleration☆19Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆422Updated this week
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆28May 31, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Libegpu is a library for enumerating eGPU devices & enclosures.☆17Nov 24, 2025Updated 9 months ago
- ☆37Jul 17, 2026Updated last month
- llama-server start/stop scripts for Qwen3.6-35B-A3B UD-Q8_K_XL GGUF on DGX Spark☆27Jul 6, 2026Updated last month
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆101Jul 11, 2026Updated last month
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆43Jun 28, 2026Updated last month
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆79Jul 18, 2026Updated last month
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆33Aug 17, 2026Updated last week
- ☆17Jun 27, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Run vLLM on 1-to-N NVIDIA DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs☆127Jun 22, 2026Updated 2 months ago
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆451Jul 3, 2026Updated last month
- Pure Rust Inference Engine☆659Updated this week
- Solana's unofficial Elixir SPL interface☆15Mar 23, 2022Updated 4 years ago
- Examples using MLX Swift☆13Apr 9, 2025Updated last year
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated last month
- Utility library for managing keys and signing with Apple's Keychain and Secure Enclave☆14Mar 15, 2022Updated 4 years ago
- Bloom Filter implementation in pure Elixir☆15Jan 8, 2021Updated 5 years ago
- Merge & adjust volume of multiple audios into video☆12Nov 8, 2020Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆31Mar 2, 2026Updated 5 months ago
- 🏛️ Hermes Gate — Terminal TUI for managing remote Hermes Agent sessions with auto-reconnect, detach support, and zero config☆26Apr 21, 2026Updated 4 months ago
- iOS audio recording tool for quickly creating song demos.☆15Aug 15, 2019Updated 7 years ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,159Updated this week
- Let any AI agent run DaVinci Resolve — MCP server with live Studio control, free-edition FCPXML/LUT workflows, reference color matching, …☆31Jul 20, 2026Updated last month
- A simple, cross-platform CLI tool for quickly switching between Claude Code configuration profiles by managing different settings.json ve…☆17Jan 17, 2026Updated 7 months ago
- ITK IO for images stored in OME-Zarr format.☆11Sep 4, 2025Updated 11 months ago