LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
☆62Jul 11, 2026Updated 3 weeks ago
Alternatives and similar repositories for llm-inference-bench
Users that are interested in llm-inference-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆154Updated this week
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆51Updated this week
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 5 months ago
- ☆16Jan 15, 2026Updated 6 months ago
- These are performance benchmarks we did to prepare for our own privacy-preserving and NDA-compliant in-house AI coding assistant. If by a…☆32Apr 2, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Your AI colleague, in the apps you already use.☆15Updated this week
- Build different debian suites for the radxa rock 4 se☆12Mar 11, 2024Updated 2 years ago
- NVIDIA Linux open GPU with P2P support☆377Jul 13, 2026Updated 3 weeks ago
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆75Jul 22, 2026Updated last week
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆466Apr 17, 2026Updated 3 months ago
- DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB)☆21Jun 9, 2026Updated last month
- ☆34Jul 19, 2026Updated 2 weeks ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆38Jul 9, 2026Updated 3 weeks ago
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆17Jul 3, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆16May 11, 2026Updated 2 months ago
- ☆15May 23, 2026Updated 2 months ago
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆21Apr 28, 2026Updated 3 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,795Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆2,990Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆22Apr 16, 2026Updated 3 months ago
- Agent-friendly GPU profile-query CLI☆107Updated this week
- ☆25Oct 22, 2025Updated 9 months ago
- Python class for the Ender 3 V2 LCD☆26Jan 6, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆11Jun 21, 2023Updated 3 years ago
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆38Jul 13, 2026Updated 3 weeks ago
- Catches silent CPU fallback and mislabeled tok/s in local LLMs. llama.cpp and ollama, one file, no deps.☆17Jul 16, 2026Updated 2 weeks ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆171Updated this week
- ASAP smoothing☆13Sep 8, 2017Updated 8 years ago
- A Go library for working with rows, columns, or matrix (deprecated, see https://github.com/shuLhan/share/tree/master/lib/tabula).☆11Nov 28, 2018Updated 7 years ago
- Ship your repo + live coding-agent session (Claude Code / Codex / pi / Droid) to another machine over Tailscale; it resumes in tmux and k…☆170Jul 4, 2026Updated last month
- High-Resolution Differential Z-Belt Mod for V0 (with optional Kirigami support)☆12May 22, 2022Updated 4 years ago
- ☆10Jan 22, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Release repository for intric☆21Feb 20, 2026Updated 5 months ago
- Explore solana accounts via an interconnected graph interface☆15Jul 7, 2026Updated 3 weeks ago
- Gemma 4 local multimodal lab with React UI, FastAPI backend, TTS, and RTX 5090 benchmark reports.☆19Apr 16, 2026Updated 3 months ago
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆21Jul 8, 2026Updated 3 weeks ago
- A distributed work-item database for agent swarms, using git as the sync layer☆24Updated this week
- Encapsulate dom-anchor-text-quote and dom-anchor-text-position for use in browser scripts☆14Sep 2, 2021Updated 4 years ago
- This is a repo to compile all the 3d printer modifications☆11Sep 24, 2022Updated 3 years ago