A Practitioner handbook for production llm serving.
☆189Sep 5, 2026Updated last week
Alternatives and similar repositories for llm-inference-at-scale
Users that are interested in llm-inference-at-scale are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A physics-grounded, cost-aware optimization loop for vLLM☆66Aug 22, 2026Updated 3 weeks ago
- Dasein managed vector index service - Python SDK☆23May 11, 2026Updated 4 months ago
- An Efficient Compression Framework for LLM☆66Jun 27, 2026Updated 2 months ago
- An open-sourced Pokemon Go battle simulator☆18Dec 18, 2022Updated 3 years ago
- ☆10Jul 20, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Frontier-grade answers from any mix of models — a local MCP server bringing OpenRouter's Fusion panel architecture to any MCP client. Bri…☆34Updated this week
- Hands-on Kubernetes troubleshooting lab with 15 production-grade scenarios covering workloads, networking, storage, security, autoscaling…☆27Jul 21, 2026Updated last month
- Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.☆18Aug 14, 2026Updated 3 weeks ago
- CDRAG is a new retrieval framework that uses hierarchical document clustering, and LLM-guided document selection from those clusters to c…☆38Apr 13, 2026Updated 5 months ago
- I developed a fine-tuned retrieval head for RAG that learns to more reliably retrieve relevant passages by transforming the query embeddi…☆17May 21, 2026Updated 3 months ago
- A simple but instructive implementation of DP, TP, FSDP, FSDP+TP using pytorch distributed primitives☆24Apr 12, 2026Updated 5 months ago
- One script to replace every outdated tool that Apple ships on macOS with modern versions via Homebrew☆17Jul 18, 2026Updated last month
- ☆25Dec 15, 2025Updated 8 months ago
- The code implementation of HyGRAG, accepted by WWW'26.☆15May 31, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated 3 months ago
- GPU Engineering for AI Systems☆596May 17, 2026Updated 3 months ago
- ☆13Dec 12, 2021Updated 4 years ago
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆75Jul 6, 2026Updated 2 months ago
- ☆19Updated this week
- Programmable automated machine learning - proof of concept☆15Oct 9, 2024Updated last year
- Data Science Course☆15Jan 8, 2024Updated 2 years ago
- The rag pipeline for optimizing dynamic data editing.☆23Oct 30, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Demo of fine-tuning QA models for answering FAQ of cloud providers documentation☆11Jun 20, 2026Updated 2 months ago
- This document aims to agree on a broad, international strategy for the implementation of open scholarship that meets the needs of differe…☆11Dec 10, 2018Updated 7 years ago
- ☆21Jan 22, 2026Updated 7 months ago
- Matra-Hachette Alice MC-10 for MiSTer FPGA☆12Dec 8, 2025Updated 9 months ago
- Open Source client side code of INDRZ indoor wayfinding, mapping, routing system. Mirrored from https://gitlab.com/indrz.☆14Feb 16, 2025Updated last year
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆12Nov 26, 2024Updated last year
- Multi-strategy RAG system achieving 74% Recall@10 on MultiHop-RAG. Combines RAPTOR hierarchical retrieval, knowledge graphs, HyDE, BM25, …☆42Feb 3, 2026Updated 7 months ago
- A simple SPA to implement User-Owns-Data Embedding in Salesforce.☆12Aug 17, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Streaming Retrieval-Augmented Generation (RAG) agent in Go. It consumes real-time data from Kafka topics, processes it in configurable wi…☆27Jun 7, 2025Updated last year
- Bulk Salesforce Files downloader built with Robot Framework and Python. Supports parallel batches, validated downloads, CI, and Data Load…☆24Aug 27, 2026Updated 2 weeks ago
- Reliable RAG setup that uses Semantic Double Merging Chunking from llamaindex, Qdrant Hybrid Search, colBERT for reranking and Google Gem…☆41Dec 15, 2024Updated last year
- rudradb-opin-examples is for example implementations of the pip install rudradb-opin☆29Mar 3, 2026Updated 6 months ago
- ☆10Feb 12, 2024Updated 2 years ago
- [ACL 2026] WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora☆18May 11, 2026Updated 4 months ago
- The second generation of the Triangle Regional Model☆13Aug 27, 2026Updated 2 weeks ago