LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
☆107Oct 2, 2026Updated this week
Alternatives and similar repositories for llm-inference-bench
Users that are interested in llm-inference-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆270Updated this week
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆83Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,117Updated this week
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 7 months ago
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆17Jan 15, 2026Updated 8 months ago
- A docs-first guide to LLM system design — hybrid search, embedding pipelines, reranking, and LLM-as-judge patterns.☆47Jun 22, 2026Updated 3 months ago
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆51Aug 17, 2026Updated last month
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆47Updated this week
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆77Jul 11, 2026Updated 2 months ago
- These are performance benchmarks we did to prepare for our own privacy-preserving and NDA-compliant in-house AI coding assistant. If by a…☆32Apr 2, 2025Updated last year
- Your AI colleague, in the apps you already use.☆14Updated this week
- Claude Code conversation history☆24Jul 31, 2025Updated last year
- ☆12Oct 9, 2023Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Build different debian suites for the radxa rock 4 se☆12Mar 11, 2024Updated 2 years ago
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆83Jul 22, 2026Updated 2 months ago
- Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"☆11Mar 31, 2024Updated 2 years ago
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆143Sep 12, 2026Updated 3 weeks ago
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆25Nov 26, 2025Updated 10 months ago
- ☆35Updated this week
- This project is a modernized reimplementation of the original Speck molecule renderer created by Rye Terrell.☆21Feb 22, 2026Updated 7 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,272Updated this week
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TokenSpeed is a speed-of-light LLM inference engine.☆2,189Updated this week
- Python class for the Ender 3 V2 LCD☆26Jan 6, 2023Updated 3 years ago
- Agent-friendly GPU profile-query CLI☆127Aug 21, 2026Updated last month
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆29Sep 3, 2026Updated last month
- ☆11Jun 21, 2023Updated 3 years ago
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆20Updated this week
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated 2 months ago
- Small program to run requests against a web server and look for problems☆11Jan 20, 2016Updated 10 years ago
- ASAP smoothing☆13Sep 8, 2017Updated 9 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Generate AI Art with Fooocus via a Discord Bot☆19Sep 8, 2026Updated 3 weeks ago
- Verified control of the whole desktop for agents on Windows and Linux: apps with no API, windows, shell, browser. MCP server + CLI + REST…☆27Sep 6, 2026Updated 3 weeks ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆228Sep 21, 2026Updated last week
- High-Resolution Differential Z-Belt Mod for V0 (with optional Kirigami support)☆12May 22, 2022Updated 4 years ago
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆19Jul 13, 2026Updated 2 months ago
- Your shipyard for parallel development workflows. Maintain your digital yard with clean branches, productive workflows, and AI-era readin…☆19Jul 23, 2026Updated 2 months ago
- ☆53Updated this week