LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
☆92Sep 6, 2026Updated this week
Alternatives and similar repositories for llm-inference-bench
Users that are interested in llm-inference-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆213Updated this week
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆77Updated this week
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,038Updated this week
- ☆37Apr 26, 2026Updated 4 months ago
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 7 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago
- ☆23Aug 29, 2026Updated 2 weeks ago
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆26Jul 23, 2026Updated last month
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆49Aug 17, 2026Updated 3 weeks ago
- A simple mouse reporter☆19Apr 18, 2022Updated 4 years ago
- These are performance benchmarks we did to prepare for our own privacy-preserving and NDA-compliant in-house AI coding assistant. If by a…☆32Apr 2, 2025Updated last year
- ☆12Sep 9, 2024Updated 2 years ago
- A plugin architecture for React.☆12May 10, 2021Updated 5 years ago
- Solve algorithmic Python challenges to sharpen the tools.☆11Apr 13, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Your AI colleague, in the apps you already use.☆14Updated this week
- Claude Code conversation history☆24Jul 31, 2025Updated last year
- ☆11Jul 24, 2025Updated last year
- Build different debian suites for the radxa rock 4 se☆12Mar 11, 2024Updated 2 years ago
- NVIDIA Linux open GPU with P2P support☆511Sep 4, 2026Updated last week
- ☆12Aug 1, 2016Updated 10 years ago
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆83Jul 22, 2026Updated last month
- Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"☆11Mar 31, 2024Updated 2 years ago
- Generate cordova/splash files from a single svg, and update config.xml☆12Feb 25, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Emulation Station for Windows 10☆11Apr 26, 2021Updated 5 years ago
- Embedded sandbox with MCP proxy☆36Apr 15, 2026Updated 4 months ago
- 📡 GPS satellite tracker - which GPS satellites are currently visible above a given location on earth?☆17Nov 6, 2019Updated 6 years ago
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆340Updated this week
- punt makes adb logcat better☆15Apr 28, 2022Updated 4 years ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆17Updated this week
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆24Nov 26, 2025Updated 9 months ago
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆504Apr 17, 2026Updated 4 months ago
- Fully local code indexing and sematic search tool☆16May 20, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ⚛️ 🖨️ A hierarchy sensitive, middleware-defined console.log for React and React Native. ✨☆21Nov 4, 2020Updated 5 years ago
- Transform coordinates from one coordinate system to another.☆19Nov 6, 2023Updated 2 years ago
- a simple API of disc golf disc profiles☆13Aug 19, 2022Updated 4 years ago
- A <HeatMap /> Native Module component for React Native.☆19Jan 4, 2023Updated 3 years ago
- 🤯💬 A Turing-Complete Binary Brainfuck Interpreter built only using CSS & HTML☆17May 15, 2024Updated 2 years ago
- 🕊️ Is it a bird? ✈️ Is it a plane? No, it's a search-bar-button! ⚛️☆22Sep 27, 2020Updated 5 years ago
- This project is a modernized reimplementation of the original Speck molecule renderer created by Rye Terrell.☆21Feb 22, 2026Updated 6 months ago