Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.
☆40Sep 28, 2026Updated this week
Alternatives and similar repositories for llm-bench-rig
Users that are interested in llm-bench-rig are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Meta-research lab for domain chip R&D, methodology improvement, and recursive self-improvement☆22Jul 26, 2026Updated 2 months ago
- Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for no…☆52Sep 20, 2026Updated 2 weeks ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆81Jul 18, 2026Updated 2 months ago
- An Enhanced TOP program to monitor your Nvidia DGX SPARK's Hardware☆37Jan 6, 2026Updated 8 months ago
- Docker image description with the newest Renode version☆15Sep 7, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆133Updated this week
- support kubernetes feature for autogen(https://github.com/microsoft/autogen)☆11Sep 15, 2025Updated last year
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- An llm wrapper for OpenAI☆13Dec 14, 2024Updated last year
- ☆10Jul 14, 2025Updated last year
- Drop-in prompt-caching fixes for the LLM agent harness you use. Point your AI coding agent at this repo and it ships the patches.☆113Aug 29, 2026Updated last month
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆27Jul 23, 2026Updated 2 months ago
- ☆10Dec 19, 2024Updated last year
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆86Jun 28, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An experiment to see if chatgpt can improve the output of the stanford alpaca dataset☆12Mar 29, 2023Updated 3 years ago
- Self Sovereign Creative AI Operating System and Task Execution Engine☆19Jun 11, 2026Updated 3 months ago
- A curated collection of production-ready Hermes Agent skills — brainstorming, PRD workflows, debugging, Apple integrations, MLOps, docume…☆78May 2, 2026Updated 5 months ago
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 3 months ago
- Hermes Relay for Chrome gives Hermes Agent a direct browser surface for page context, capture, watchlists, and AI handoff.☆23May 4, 2026Updated 5 months ago
- pastescript template for a complete django project with pip+virtualenv, fabric, a gevent-based wsgi server and various helpers scripts. R…☆13Sep 29, 2010Updated 16 years ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 3 months ago
- small board to convert linear 3 pin regulator to switching regulator (TO220 compatible, eg 7805)☆11Oct 8, 2015Updated 10 years ago
- A set of tools for building AI coders