Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
☆76Jul 25, 2026Updated last month
Alternatives and similar repositories for GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
Users that are interested in GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- ☆17May 10, 2026Updated 3 months ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆81Jul 13, 2026Updated last month
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆38Jul 13, 2026Updated last month
- Self-contained DeepSeek V4 Flash DSpark TP=2 recipe for 2x DGX Spark with 62 tok/s benchmark☆24Jul 13, 2026Updated last month
- Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.☆80Jun 29, 2026Updated 2 months ago
- Persistent JavaScript REPL for context manipulation, MCP integration, and focused sub-agent orchestration. A pi extension.☆25May 30, 2026Updated 3 months ago
- ☆16Feb 21, 2026Updated 6 months ago
- A self-evolving agent in rust with one tool that only uses skills☆31Aug 12, 2026Updated 3 weeks ago
- Using Statuscake's monitoring and Cloudflare to achieve high availability.☆24May 27, 2015Updated 11 years ago
- Open Source Desktop App to help with Previs for AI-native filmmaking — stage grey-box scenes, choreograph camera & cast with marks, expor…☆117Aug 5, 2026Updated 3 weeks ago
- ☆88Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆80Jun 28, 2026Updated 2 months ago
- Homebridge plugin for Pentair IntelliCenter☆12Apr 13, 2023Updated 3 years ago
- my version of the models dev site which is way faster and has the filters I want☆18Apr 22, 2026Updated 4 months ago
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆138Aug 26, 2026Updated last week
- ☆12Feb 19, 2017Updated 9 years ago
- ☆26Apr 23, 2023Updated 3 years ago
- Interface boards for the Neato LDS - the Lidar sensor of the XV-11.☆14Jun 21, 2014Updated 12 years ago
- ☆16Jun 25, 2026Updated 2 months ago
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆103Jul 11, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- Pure Rust Inference Engine☆675Updated this week
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,749Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆455Jul 3, 2026Updated 2 months ago
- Cut your LLM output tokens up to 76%. System prompt rules for any model. No install required.☆23Apr 5, 2026Updated 4 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- ☆17Sep 5, 2021Updated 4 years ago
- SciKit-Learn compatible library for training mixed-effects models.☆13Jun 9, 2023Updated 3 years ago
- Sensirion SVM30 on Arduino, ESPxx, 32U4, Lorawan, DUE☆10Oct 30, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Serial buffer for Particle Photon and Electron☆13Nov 2, 2024Updated last year
- An open-source hardware racing drone design and demonstration RL software for the artificial intelligence robotic drone race competitions☆13Sep 6, 2020Updated 5 years ago
- NLSpec instruction following benchmark for https://factory.strongdm.ai/products/attractor☆21Feb 26, 2026Updated 6 months ago
- Neovim plugin that renders Mermaid diagrams as inline ASCII art using virtual text☆38Jul 11, 2026Updated last month
- ☆16May 21, 2026Updated 3 months ago
- AseeVR fix for Droolon Pi1☆14Mar 16, 2024Updated 2 years ago
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 4 months ago