Multiple NVIDIA GPUs or Apple Silicon for Large Language Model Inference?
☆1,939May 13, 2024Updated 2 years ago
Alternatives and similar repositories for GPU-Benchmarks-on-LLM-Inference
Users that are interested in GPU-Benchmarks-on-LLM-Inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A fast inference library for running LLMs locally on modern consumer-class GPUs☆4,631Mar 4, 2026Updated 7 months ago
- LLM inference in C/C++☆130,372Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆93,218Updated this week
- Large-scale LLM inference engine☆1,868Sep 11, 2026Updated 3 weeks ago
- Web UI for ExLlamaV2☆517Feb 5, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Python bindings for llama.cpp☆10,639Updated this week
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.☆77,224Updated this week
- High-speed Large Language Model Serving for Local Deployment☆9,817May 11, 2026Updated 4 months ago
- Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.☆47,723Aug 17, 2026Updated last month
- Run frontier AI locally.☆47,748Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,795Updated this week
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.