LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
☆76Aug 14, 2026Updated last week
Alternatives and similar repositories for llm-inference-bench
Users that are interested in llm-inference-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆184Updated this week
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆66Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆874Updated this week
- ☆37Apr 26, 2026Updated 3 months ago
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆25Jul 23, 2026Updated last month
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆46Updated this week
- These are performance benchmarks we did to prepare for our own privacy-preserving and NDA-compliant in-house AI coding assistant. If by a…☆32Apr 2, 2025Updated last year
- NVIDIA Linux open GPU with P2P support☆441Updated this week
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆24Nov 26, 2025Updated 8 months ago
- A simple replacement for std::unordered_map☆51Jul 27, 2024Updated 2 years ago
- DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB)☆22Jun 9, 2026Updated 2 months ago
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆18Jul 3, 2026Updated last month
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- BabyCube: small linear rails coreXY 3D printer☆69Nov 6, 2024Updated last year
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆24Apr 28, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,077Aug 15, 2026Updated last week
- Agent-friendly GPU profile-query CLI☆118Updated this week
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆21Aug 17, 2026Updated last week
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated last month
- Catch local LLMs spilling into the CPU. One run shows GPU placement, prefill, decode, memory, power and a verdict.☆17Updated this week
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆177Aug 9, 2026Updated 2 weeks ago
- High-Resolution Differential Z-Belt Mod for V0 (with optional Kirigami support)☆12May 22, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆53Updated this week
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆20Jul 13, 2026Updated last month
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆21Jul 11, 2026Updated last month
- A small operating file for keeping long-running AI agent sessions from degrading. Capable models can do useful long work. The problem is…☆30Jun 18, 2026Updated 2 months ago
- Explore solana accounts via an interconnected graph interface☆15Jul 7, 2026Updated last month
- Agentic SDLC with local LLM's.☆17Updated this week
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆33Aug 17, 2026Updated last week
- Encapsulate dom-anchor-text-quote and dom-anchor-text-position for use in browser scripts☆14Sep 2, 2021Updated 4 years ago
- This is a repo to compile all the 3d printer modifications☆11Sep 24, 2022Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Run LLMs with MLX☆15May 1, 2026Updated 3 months ago
- ☆11Nov 29, 2017Updated 8 years ago
- A small test for multithreaded C++ stack unwinding on unixes☆16Feb 24, 2020Updated 6 years ago
- Bilinear Pairings Components Library for Delphi☆12Dec 19, 2018Updated 7 years ago
- A mount for a standard 3x15mm cartridge thermistor on 1515 T-Slot Extrusion☆11Feb 23, 2023Updated 3 years ago
- Example of converting a module to a web-worker in electron☆13Jan 25, 2018Updated 8 years ago
- Klipper plugin to make system IP/etc available to macros☆68Aug 2, 2023Updated 3 years ago