Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch, per-GPU metrics, shareable summary. Zero deps.
☆26Jul 23, 2026Updated last month
Alternatives and similar repositories for 2Wild-Coding-Agent-Latency-Monitor
Users that are interested in 2Wild-Coding-Agent-Latency-Monitor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆61Jul 30, 2026Updated last month
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆105Sep 5, 2026Updated last week
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆400Updated this week
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆48Aug 17, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated 3 weeks ago
- Sglang LLM Inference: RTX Pro 6000 vs DGX Spark☆22Oct 18, 2025Updated 10 months ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆106Aug 20, 2026Updated 3 weeks ago
- ☆17May 10, 2026Updated 4 months ago
- Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audi…☆25Aug 16, 2026Updated 3 weeks ago
- ☆40Jul 17, 2026Updated last month
- ☆15Aug 16, 2023Updated 3 years ago
- Script Center for System Center Configuration Manager☆13Jul 20, 2023Updated 3 years ago
- DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks☆1,350Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Rust CLI tool that transforms your project files into perfectly formatted context blocks for Large Language Models. Ideal for code revi…☆13May 11, 2025Updated last year
- AI Liquidity Management Agent☆13Jan 19, 2026Updated 7 months ago
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆92Updated this week
- Mixed-vendor GPU inference cluster manager with speculative decoding☆34Jul 2, 2026Updated 2 months ago
- A "polyfill" that standardizes and centralizes AGENTS.md configuration and Agent Skills support within your repository.☆19Feb 19, 2026Updated 6 months ago
- Pure Rust Inference Engine☆679Updated this week
- React hook for Google Gemini Live API - real-time voice streaming, session recording, workflow automation, and smart element detection☆26Jan 20, 2026Updated 7 months ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆541Updated this week
- Tools to speed up migration from other SSE solutions to Microsoft's Global Secure Access☆19Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🐣🕐📅 A simple utility to draft scheduling emails.☆12Sep 13, 2023Updated 2 years ago
- Radiantloom Email Assist 7B is an email-assistant large language model fine-tuned from Zephyr-7B-Beta, over a custom-curated dataset of 1…☆14Jan 19, 2024Updated 2 years ago
- ☆19Feb 24, 2025Updated last year
- Prompts for Claude Code to use with Moonshot's Kimi K2 Thinking Model☆15Nov 12, 2025Updated 10 months ago
- Unofficial BambuStudio tools and scripts☆37Jan 5, 2025Updated last year
- Generic sharded thread safe LRU cache in Go.☆12Apr 13, 2022Updated 4 years ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆82Jul 13, 2026Updated 2 months ago
- SLA Buddy: a helpful robot to help you meet Service Level Agreement in Slack☆10Apr 5, 2024Updated 2 years ago
- Versatile Metrics Collection for Python☆19Jun 15, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- deep seek & o1 auto coders which write python code from a simple description and iteratively improvesit and fix errors☆99Jan 21, 2025Updated last year
- The Package Support Framework (PSF) is a kit for applying compatibility fixes to packaged desktop applications.☆24Jul 2, 2026Updated 2 months ago
- TQ3 KV cache compression for ComfyUI. 4.6x VRAM savings for video generation. Enables LTX-2.3 22B on V100 32GB.☆42Aug 20, 2026Updated 3 weeks ago
- Sentence Embedding as a Service☆15Jun 30, 2025Updated last year
- PowerShell Module to find compatible FIDO2 keys for Entra☆20Updated this week
- Device Serial Number Import Tool for Intune Autopilot V2☆17Apr 23, 2026Updated 4 months ago
- Framework to achieve context distillation in LLMs☆15Nov 24, 2023Updated 2 years ago