Everything you need to know about LLM inference
☆440Sep 23, 2026Updated this week
Alternatives and similar repositories for llm-inference-handbook
Users that are interested in llm-inference-handbook are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- API serving for your diffusers models☆11Jan 19, 2024Updated 2 years ago
- Benchmark and optimize LLM inference across frameworks with ease☆201Jul 14, 2026Updated 2 months ago
- Simple dependency injection framework for Python☆21Jul 14, 2026Updated 2 months ago
- ☆15Jul 5, 2025Updated last year
- ☆24Sep 1, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆20Mar 20, 2025Updated last year
- learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen☆4,706Updated this week
- Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, T…☆511Updated this week
- ☆16Oct 15, 2025Updated 11 months ago
- Command-line tool for debugging MCP servers☆37Updated this week
- ☆18Apr 28, 2025Updated last year
- Postgres extension that speeds up analytics queries by upto 90%☆52Jun 8, 2024Updated 2 years ago
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.☆18Updated this week
- A concise, beginner-friendly introduction to the core ideas of linear algebra.☆2,001Mar 16, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆484Nov 25, 2025Updated 9 months ago
- A reverse literate programming tool.☆16Jun 26, 2024Updated 2 years ago
- A domain specific language to build complex workflows☆15Jun 26, 2026Updated 2 months ago
- This is a python implementation for stitching images.☆229Oct 3, 2024Updated last year
- libcmods - provides c module headers for popular libraries☆15Oct 12, 2014Updated 11 years ago
- Terminal user interface for a Kanban board☆11Nov 5, 2021Updated 4 years ago
- Online Inference API for NLP Transformer models - summarization, text classification, sentiment analysis and more☆45Mar 16, 2024Updated 2 years ago
- Provides deploy scripts and CSI for Lustre.☆14Apr 13, 2026Updated 5 months ago
- ☆16May 23, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆18Sep 17, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMs☆92,527Updated this week
- A blueprint for AI development, focusing on applied examples of RAG, information extraction, analysis and fine-tuning in the age of LLMs …☆67Feb 6, 2025Updated last year
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual h…☆201Updated this week
- Manage kubernetes node-level kernel tuning ( using sysctl ).☆30Nov 21, 2025Updated 10 months ago
- ☆168Mar 5, 2026Updated 6 months ago
- Simple Agents Made Easy☆617Mar 16, 2026Updated 6 months ago
- A comprehensive cookbook demonstrating how to implement CrewAI with Anthropic's prompt caching feature for efficient LLM operations☆16Aug 11, 2025Updated last year
- A demonstration of text/GUI bi-directional editing via an LSP server☆38Jul 1, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆17Oct 21, 2025Updated 11 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆850Updated this week
- ☆12Jan 11, 2024Updated 2 years ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,895Updated this week
- Nano vLLM☆15,600Apr 26, 2026Updated 4 months ago
- Build and Deploy a voice-based Chatbot with Langchain and BentoML☆18May 2, 2023Updated 3 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,377Updated this week