Interactive launcher and benchmarking harness for llama.cpp server throughput, with tests, sweeps, and round‑robin load tools.
☆441Feb 8, 2026Updated 5 months ago
Alternatives and similar repositories for llama-throughput-lab
Users that are interested in llama-throughput-lab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmark tool for measuring speculative decoding speedups. Sweep draft/target model combinations and generate interactive charts.☆171Feb 18, 2026Updated 5 months ago
- ☆303May 18, 2026Updated 2 months ago
- llama-benchy - llama-bench style benchmarking tool for all backends☆613Jul 10, 2026Updated 3 weeks ago
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated 3 weeks ago
- AutoBE-generated backend application examples☆24Jun 30, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,251Updated this week
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆126May 27, 2026Updated 2 months ago
- Interested in running a conspiracy of agents (technical term) on your Local AI Infra? Who isn't!☆16Mar 30, 2026Updated 4 months ago
- ☆487Updated this week
- ☆41Sep 22, 2025Updated 10 months ago
- This is a FastAPI based LLM server. Load multiple LLM models (MLX or llama.cpp) simultaneously using multiprocessing.☆18Apr 8, 2026Updated 3 months ago
- ☆17Jul 21, 2026Updated last week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,209Updated this week
- Docker configuration for running VLLM on dual DGX Sparks☆1,944Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated last month
- LLM inference in C/C++☆122,564Updated this week
- Run your own AI cluster at home with everyday devices 📱💻 🖥️⌚☆22Jul 14, 2025Updated last year
- Don't bug your friends with articles they'll never read. AI's have infinite attention, leverage them instead! Use the curation buddy to e…☆21May 2, 2024Updated 2 years ago
- Tool to explore an Appwrite project☆19Mar 19, 2026Updated 4 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 3 months ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- Distributed packaging system for Arch Linux with Github Actions☆17Updated this week
- Suckless no dependencies CUDA/C++ port of NVIDIA's DVLT, feed it images, get a 3D point cloud + camera poses. NO python. One fast 5MB bin…☆60Jun 4, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆98Mar 8, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,990Updated this week
- LLM inference in C/C++☆64May 7, 2026Updated 2 months ago
- ☆176Apr 7, 2026Updated 3 months ago
- One-click LLM server with TurboQuant Llama CPP engine☆20Apr 16, 2026Updated 3 months ago
- Run Claude Code (and codex) to generate a project plan, then run them in a loop for days until they're done☆14Jan 18, 2026Updated 6 months ago
- ☆250Oct 30, 2025Updated 9 months ago
- Diffusion_TTS extension for booga☆71Sep 6, 2025Updated 10 months ago
- OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous bat…☆1,485Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A quick and optimized solution to manage llama based gguf quantized models, download gguf files, retreive messege formatting, add more mo…☆13Jan 13, 2024Updated 2 years ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆34Jul 22, 2026Updated last week
- ☆20Jan 3, 2026Updated 7 months ago
- Tutorial Assets☆619Dec 15, 2025Updated 7 months ago
- ☆20May 30, 2025Updated last year
- Host your own dyndns server with docker!☆18Dec 2, 2020Updated 5 years ago
- ☆15Apr 3, 2025Updated last year