Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
☆73Jul 25, 2026Updated 2 weeks ago
Alternatives and similar repositories for GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
Users that are interested in GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 655,360-token context + MTP spec decode (DCP4, fp8_ds_mla) on a 4x NVIDIA DGX Spark …☆29Jul 13, 2026Updated last month
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆18Jul 5, 2026Updated last month
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆626Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆363Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆98Jul 9, 2026Updated last month
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆36Jul 12, 2026Updated last month
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Jul 7, 2026Updated last month
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆213Updated this week
- Self-contained DeepSeek V4 Flash DSpark TP=2 recipe for 2x DGX Spark with 62 tok/s benchmark☆24Jul 13, 2026Updated last month
- Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.☆73Jun 29, 2026Updated last month
- Persistent JavaScript REPL for context manipulation, MCP integration, and focused sub-agent orchestration. A pi extension.☆24May 30, 2026Updated 2 months ago
- ☆36Jul 17, 2026Updated 3 weeks ago
- ☆16Feb 21, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A self-evolving agent in rust with one tool that only uses skills☆22Jul 15, 2026Updated 3 weeks ago
- Wasserman's Filmmaker Suite — ScriptBreak, Cork Board, Master Canvas, Blockout, Motion Previs Studio, Storyboard Reference Studio, Circle…☆105Updated this week
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆61Jul 30, 2026Updated last week
- GPT-4o Powered Calorie Detecor☆18May 29, 2024Updated 2 years ago
- Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)☆34Jul 6, 2026Updated last month
- ☆19Oct 11, 2025Updated 10 months ago
- Open Source Desktop App to help with Previs for AI-native filmmaking — stage grey-box scenes, choreograph camera & cast with marks, expor…☆106Aug 5, 2026Updated last week
- ☆84Updated this week
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆79Jun 28, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆45Nov 20, 2025Updated 8 months ago
- Output high level Pcode (PcodeAST) in Ghidra☆17Apr 7, 2023Updated 3 years ago
- OpenCodeReview (Community Edition) - Agentic AI based source code review tool powered by NVIDIA NeMo Agent Toolkit. NVIDIA Hackathon Awar…☆19Sep 27, 2025Updated 10 months ago
- Homebridge plugin for Pentair IntelliCenter☆12Apr 13, 2023Updated 3 years ago
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆26Jul 8, 2026Updated last month
- This repo helps to transform text into a better form for lora training☆12Apr 9, 2023Updated 3 years ago
- ☆18Jul 9, 2026Updated last month
- Guidelines for using Cython (documentation)☆12Mar 4, 2023Updated 3 years ago
- Interface boards for the Neato LDS - the Lidar sensor of the XV-11.☆14Jun 21, 2014Updated 12 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A easy to use CLI to generate JSONL datasets from a TXT file using LLMs.☆25Apr 24, 2026Updated 3 months ago
- ☆16Jun 25, 2026Updated last month
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆101Jul 11, 2026Updated last month
- Run Geoffrey Huntley's Ralph Wiggum loop inside Pi, no shell scripting needed. Native commands control the loop, and the bundled agent sk…☆25Aug 5, 2026Updated last week
- ITK IO for images stored in OME-Zarr format.☆11Sep 4, 2025Updated 11 months ago
- ☆12Dec 16, 2024Updated last year
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year