Interactive launcher and benchmarking harness for llama.cpp server throughput, with tests, sweeps, and round‑robin load tools.
☆457Feb 8, 2026Updated 6 months ago
Alternatives and similar repositories for llama-throughput-lab
Users that are interested in llama-throughput-lab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmark tool for measuring speculative decoding speedups. Sweep draft/target model combinations and generate interactive charts.☆172Feb 18, 2026Updated 6 months ago
- ☆317Apr 13, 2026Updated 4 months ago
- ☆226Feb 1, 2026Updated 6 months ago
- ☆304Updated this week
- llama-benchy - llama-bench style benchmarking tool for all backends☆667Jul 10, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆1,872Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,449Updated this week
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆127May 27, 2026Updated 2 months ago
- Official implementation of OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on☆13Feb 27, 2024Updated 2 years ago
- Interested in running a conspiracy of agents (technical term) on your Local AI Infra? Who isn't!☆16Mar 30, 2026Updated 4 months ago
- ☆41Sep 22, 2025Updated 11 months ago
- This is a FastAPI based LLM server. Load multiple LLM models (MLX or llama.cpp) simultaneously using multiprocessing.☆18Apr 8, 2026Updated 4 months ago
- Universal text classifier for generative models☆25Jul 25, 2024Updated 2 years ago
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,445Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Docker configuration for running VLLM on dual DGX Sparks☆2,159Updated this week
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated 2 months ago
- Don't bug your friends with articles they'll never read. AI's have infinite attention, leverage them instead! Use the curation buddy to e…☆21May 2, 2024Updated 2 years ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 4 months ago
- htmxjs is the server-side js framework for HTMX☆20Oct 22, 2023Updated 2 years ago
- WIP: ALSA config or driver for Audient iD14☆11Aug 1, 2019Updated 7 years ago
- A comprehensive WebUI Toolkit for Resemble-AI's Chatterbox☆28Jun 7, 2025Updated last year
- Suckless no dependencies CUDA/C++ port of NVIDIA's DVLT, feed it images, get a 3D point cloud + camera poses. NO python. One fast 5MB bin…☆65Jun 4, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,077Aug 15, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Run frontier AI locally.☆46,990Jun 23, 2026Updated 2 months ago
- LLM inference in C/C++☆63May 7, 2026Updated 3 months ago
- ☆175Apr 7, 2026Updated 4 months ago
- Run Claude Code (and codex) to generate a project plan, then run them in a loop for days until they're done☆14Jan 18, 2026Updated 7 months ago
- LLM finetuned for generating symbolic music☆42Sep 4, 2024Updated last year
- ☆250Oct 30, 2025Updated 9 months ago
- A quick and optimized solution to manage llama based gguf quantized models, download gguf files, retreive messege formatting, add more mo…☆14Jan 13, 2024Updated 2 years ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆40Jul 22, 2026Updated last month
- ☆20Jan 3, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Agent Zero AI framework☆18,945Updated this week
- Tutorial Assets☆617Dec 15, 2025Updated 8 months ago
- Quick access to any large language model from your browser.☆10Feb 16, 2026Updated 6 months ago
- a browser gui for nvidia smi☆21Mar 17, 2025Updated last year
- The official application repository of Olares Market☆54Updated this week
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,040Apr 23, 2026Updated 4 months ago
- ☆33Sep 3, 2024Updated last year