Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch, per-GPU metrics, shareable summary. Zero deps.
☆23Jul 23, 2026Updated last week
Alternatives and similar repositories for 2Wild-Coding-Agent-Latency-Monitor
Users that are interested in 2Wild-Coding-Agent-Latency-Monitor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆72Jul 25, 2026Updated last week
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆20Jul 13, 2026Updated 3 weeks ago
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆61Updated this week
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆138Updated this week
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆105Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆20Jul 7, 2026Updated 3 weeks ago
- ☆16May 10, 2026Updated 2 months ago
- llama-server start/stop scripts for Qwen3.6-35B-A3B UD-Q8_K_XL GGUF on DGX Spark☆27Jul 6, 2026Updated 3 weeks ago
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆38Jul 13, 2026Updated 3 weeks ago
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆42Jul 13, 2026Updated 3 weeks ago
- How to create a private Epstein search engine☆17Feb 13, 2026Updated 5 months ago
- ☆15Aug 16, 2023Updated 2 years ago
- Script Center for System Center Configuration Manager☆12Jul 20, 2023Updated 3 years ago
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆24Jul 19, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Give any blind LLM eyes — on any machine — in one shot. Tiny local VLM (Qwen3.5-0.8B) as an OpenAI-compatible vision service. Mac (MLX) /…☆29Jul 7, 2026Updated 3 weeks ago
- Openclawd configuration audit☆18Feb 5, 2026Updated 5 months ago
- A Rust CLI tool that transforms your project files into perfectly formatted context blocks for Large Language Models. Ideal for code revi…☆13May 11, 2025Updated last year
- Show PIM role to solve a task - and group to activate the needed permission☆15May 22, 2025Updated last year
- Mixed-vendor GPU inference cluster manager with speculative decoding☆31Jul 2, 2026Updated last month
- ConfigMgr LogFiler Opener automates the usage of CMTrace, CMLogViewer and OneTrace for opening single or multiple ConfigMgr Client Logfil…☆13May 22, 2023Updated 3 years ago
- Pure Rust Inference Engine☆627Updated this week
- Vibe coded Chrome plugin that converts WPM in Nitrotype to TPS (tokens per second)☆18Nov 23, 2025Updated 8 months ago
- Single-file web UI for NVIDIA DGX Spark — pull Ollama models, browse and download from HuggingFace, manage LiteLLM routing, and control S…☆31May 19, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆491Updated this week
- 🐣🕐📅 A simple utility to draft scheduling emails.☆12Sep 13, 2023Updated 2 years ago
- Radiantloom Email Assist 7B is an email-assistant large language model fine-tuned from Zephyr-7B-Beta, over a custom-curated dataset of 1…☆14Jan 19, 2024Updated 2 years ago
- ☆19Feb 24, 2025Updated last year
- Prompts for Claude Code to use with Moonshot's Kimi K2 Thinking Model☆16Nov 12, 2025Updated 8 months ago
- A real estate agent CRM for managing clients' properties and profiles☆15May 6, 2024Updated 2 years ago
- Generic sharded thread safe LRU cache in Go.☆12Apr 13, 2022Updated 4 years ago
- MCP Server to make searching openrouter easy☆23Feb 28, 2026Updated 5 months ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆63Jul 13, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- SLA Buddy: a helpful robot to help you meet Service Level Agreement in Slack☆10Apr 5, 2024Updated 2 years ago
- deep seek & o1 auto coders which write python code from a simple description and iteratively improvesit and fix errors☆99Jan 21, 2025Updated last year
- The Package Support Framework (PSF) is a kit for applying compatibility fixes to packaged desktop applications.☆23Jul 2, 2026Updated last month
- PowerShell Module to find compatible FIDO2 keys for Entra☆19Updated this week
- Device Serial Number Import Tool for Intune Autopilot V2☆17Apr 23, 2026Updated 3 months ago
- The Treblle SDK the Django framework☆11Sep 10, 2025Updated 10 months ago
- ☆25Updated this week