Inference server benchmarking tool
☆173Sep 23, 2026Updated this week
Alternatives and similar repositories for inference-benchmarker
Users that are interested in inference-benchmarker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minimal implementation of a Byte Pair Encoding (BPE) tokenizer in Zig☆15Apr 7, 2025Updated last year
- ANE accelerated embedding models!☆20Dec 11, 2024Updated last year
- Build compute kernels and load them from the Hub.☆753Updated this week
- Crossword puzzles in your terminal.☆22Feb 4, 2026Updated 7 months ago
- Hugging Face Inference Toolkit used to serve transformers, sentence-transformers, and diffusers models.☆97Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A course on building Large Language Models☆20Mar 24, 2025Updated last year
- 👷 Build compute kernels☆213Apr 6, 2026Updated 5 months ago
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- Code and Data for Evaluating the Evaluators☆16Aug 20, 2025Updated last year
- ☆24Sep 17, 2026Updated last week
- A Go library for semantic caching with LRU eviction, supporting vector-based similarity search with pluggable embedding backends (local o…☆29Jul 20, 2026Updated 2 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆722Updated this week
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆729Sep 17, 2026Updated last week
- ☆17Aug 14, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Exposes a serialized machine learning model through a HTTP API.☆13Jun 11, 2023Updated 3 years ago
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- Implementation of the HuggingFace Xet Protocol.☆28Aug 13, 2026Updated last month
- ☆29Updated this week
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆333Updated this week
- Hugging Face Jobs☆20Jul 11, 2025Updated last year
- ☆25Oct 10, 2025Updated 11 months ago
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- ☆25Dec 13, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Manage scalable open LLM inference endpoints in Slurm clusters☆294Jul 11, 2024Updated 2 years ago
- Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs☆1,643Updated this week
- 🏋️ A unified multi-backend utility for benchmarking Transformers, Timm, PEFT, Diffusers and Sentence-Transformers with full support of O…☆341Updated this week
- Synthetic Data Generation with Execution-Based Verification and Grounding for LLM Training.☆24Feb 7, 2025Updated last year
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends☆2,549Updated this week
- ☆20Oct 5, 2025Updated 11 months ago
- ☆33Apr 19, 2025Updated last year
- An experimental autonomous research system that conducts comprehensive, multi-hour research sessions and produces book-length reports wit…☆51Dec 7, 2025Updated 9 months ago
- Clue inspired puzzles for testing LLM deduction abilities☆47Mar 19, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆160Sep 16, 2026Updated last week
- AEGIS-256☆11Apr 10, 2026Updated 5 months ago
- ☆16Dec 16, 2024Updated last year
- Evaluation repository of wikipedia index with Dria☆10Mar 14, 2024Updated 2 years ago
- Large Language Model Text Generation Inference☆10,885Mar 21, 2026Updated 6 months ago
- LLM inference in C/C++☆24Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆8,155Updated this week