Inference server benchmarking tool
☆168Jun 9, 2026Updated 2 months ago
Alternatives and similar repositories for inference-benchmarker
Users that are interested in inference-benchmarker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minimal implementation of a Byte Pair Encoding (BPE) tokenizer in Zig☆15Apr 7, 2025Updated last year
- ANE accelerated embedding models!☆20Dec 11, 2024Updated last year
- Build compute kernels and load them from the Hub.☆723Updated this week
- Crossword puzzles in your terminal.☆22Feb 4, 2026Updated 6 months ago
- Kubernetes autoscaler for deployments that consume queue in RMQ☆18Aug 6, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Hugging Face Inference Toolkit used to serve transformers, sentence-transformers, and diffusers models.☆96Aug 3, 2026Updated last week
- Where GPUs get cooked 👩🍳🔥☆402May 26, 2026Updated 2 months ago
- A course on building Large Language Models☆20Mar 24, 2025Updated last year
- 👷 Build compute kernels☆213Apr 6, 2026Updated 4 months ago
- Benchmark suite for LLMs from Fireworks.ai☆111Aug 6, 2026Updated last week
- Simple examples using Argilla tools to build AI☆59Nov 18, 2024Updated last year
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- CodeQUEST is a generalizable framework which leverages LLMs to iteratively evaluate and enhance code quality across multiple dimensions f…☆18Feb 11, 2026Updated 6 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆542Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆23May 26, 2026Updated 2 months ago
- A Go library for semantic caching with LRU eviction, supporting vector-based similarity search with pluggable embedding backends (local o…☆29Jul 20, 2026Updated 3 weeks ago
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆712Aug 2, 2026Updated 2 weeks ago
- ☆16Dec 16, 2025Updated 7 months ago
- TUI for LM Studio☆17Oct 19, 2025Updated 9 months ago
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- Implementation of the HuggingFace Xet Protocol.☆27Jun 30, 2026Updated last month
- ☆17Oct 21, 2025Updated 9 months ago
- ☆29May 26, 2026Updated 2 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Hugging Face Jobs☆20Jul 11, 2025Updated last year
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆321Updated this week
- Use safetensors with ONNX 🤗☆90Aug 4, 2026Updated last week
- ☆25Oct 10, 2025Updated 10 months ago
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- ☆25Dec 13, 2024Updated last year
- Manage scalable open LLM inference endpoints in Slurm clusters☆291Jul 11, 2024Updated 2 years ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,508Updated this week
- Patch management tool for git submodules☆17Jul 25, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.