Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 655,360-token context + MTP spec decode (DCP4, fp8_ds_mla) on a 4x NVIDIA DGX Spark (GB10) cluster
☆30Jul 13, 2026Updated last month
Alternatives and similar repositories for GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s
Users that are interested in GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆75Jul 25, 2026Updated last month
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated last month
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆43Jul 13, 2026Updated last month
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated last week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆21Aug 17, 2026Updated last week
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆975Updated this week
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆124Updated this week
- Self-contained DeepSeek V4 Flash DSpark TP=2 recipe for 2x DGX Spark with 62 tok/s benchmark☆26Jul 13, 2026Updated last month
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆440Updated this week
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆38Jul 13, 2026Updated last month
- Minimal Claude Code bootstrapper that installs an agent-architect for project-specific sub-agents.☆31Jul 1, 2026Updated last month
- Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.☆77Jun 29, 2026Updated last month
- Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config…☆33Jul 12, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated 2 months ago
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆60Jul 30, 2026Updated 3 weeks ago
- Output high level Pcode (PcodeAST) in Ghidra☆17Apr 7, 2023Updated 3 years ago
- This repo helps to transform text into a better form for lora training☆12Apr 9, 2023Updated 3 years ago
- Second Brain is a desktop application that acts as a personal knowledge base, using retrieval-augmented generation (RAG), multimodal AI m…☆27Jan 30, 2026Updated 6 months ago
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆157Aug 18, 2026Updated last week
- Eidos – A Self-Growing AI Agent with Long-Term Memory and Environmental Awareness☆23Jul 4, 2025Updated last year
- Fuzzing Infrastructure with k8s & cephfs☆12Jul 23, 2020Updated 6 years ago
- A collection of reusable Claude Code skills for automating complex development workflows☆21Feb 6, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆39Jul 17, 2026Updated last month
- Use borg and APFS snapshot to backup your Mac☆13Sep 18, 2019Updated 6 years ago
- Docker that queries Starlinks Dishy and serves you the response on it's own web server☆12Nov 6, 2021Updated 4 years ago
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆81Jul 22, 2026Updated last month
- Pipecat Guess Who?☆16Aug 1, 2025Updated last year
- KoboldCpp Smart Launcher with GPU Layer and Tensor Override Tuning☆30May 18, 2025Updated last year
- NLSpec instruction following benchmark for https://factory.strongdm.ai/products/attractor☆20Feb 26, 2026Updated 6 months ago
- ☆10Apr 16, 2021Updated 5 years ago
- A project I built in 2020 that is no longer maintained that is worth sharing for anyone who wants to play!☆17Jul 3, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- scripts to help beginners program in Bro☆21Aug 10, 2013Updated 13 years ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆31Mar 2, 2026Updated 5 months ago
- An open-source social media production system that turns a content calendar into on-brand copy, images, video, and client-ready delivery …☆34Aug 17, 2026Updated last week
- Small program to run requests against a web server and look for problems☆11Jan 20, 2016Updated 10 years ago
- RMUX implementations demos☆17May 19, 2026Updated 3 months ago
- ☆47Apr 29, 2026Updated 3 months ago
- ☆12Oct 28, 2023Updated 2 years ago