A Practitioner handbook for production llm serving.
☆146Jun 28, 2026Updated 3 weeks ago
Alternatives and similar repositories for llm-inference-at-scale
Users that are interested in llm-inference-at-scale are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Real-time terminal UI dashboard for monitoring vLLM☆24Jul 13, 2026Updated last week
- Developing K - a language model to generate OPENSCAD code from prompt☆19Dec 3, 2025Updated 7 months ago
- A curated list of resources for ML Systems Engineering - hardware, compilers, distributed training, inference, and production operations.☆26Jul 4, 2026Updated 2 weeks ago
- Become GPU kernel engineer step by step.☆60May 7, 2026Updated 2 months ago
- Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.☆16Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Curated collection of AI inference engineering resources — LLM serving, GPU kernels, quantization, distributed inference, and production …☆246Feb 4, 2026Updated 5 months ago
- A conversational AI system using Ollama with persistent memory capabilities. Features hybrid context management (sliding window + vector …☆18Mar 20, 2026Updated 4 months ago
- I developed a fine-tuned retrieval head for RAG that learns to more reliably retrieve relevant passages by transforming the query embeddi…☆17May 21, 2026Updated 2 months ago
- A simple but instructive implementation of DP, TP, FSDP, FSDP+TP using pytorch distributed primitives☆19Apr 12, 2026Updated 3 months ago
- A local-first, self-organizing AI RAG GRAPH knowledge system that reads, links and reasons over your documents — offline by default, clou…☆26Apr 22, 2026Updated 2 months ago
- ☆25Dec 15, 2025Updated 7 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆125Updated this week
- GPU Engineering for AI Systems☆545May 17, 2026Updated 2 months ago
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆71Jul 6, 2026Updated 2 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- End to end analytics solution combining Python ETL, Power BI visualization, and DAX-powered predictive modelling for Fantasy Premier Leag…☆15Feb 19, 2026Updated 5 months ago
- Demo of fine-tuning QA models for answering FAQ of cloud providers documentation☆11Jun 20, 2026Updated last month
- ☆22Jan 22, 2026Updated 5 months ago
- ☆32Mar 24, 2026Updated 3 months ago
- Open Source client side code of INDRZ indoor wayfinding, mapping, routing system. Mirrored from https://gitlab.com/indrz.☆14Feb 16, 2025Updated last year
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆11Nov 26, 2024Updated last year
- Multi-strategy RAG system achieving 74% Recall@10 on MultiHop-RAG. Combines RAPTOR hierarchical retrieval, knowledge graphs, HyDE, BM25, …☆41Feb 3, 2026Updated 5 months ago
- A simple SPA to implement User-Owns-Data Embedding in Salesforce.☆12Aug 17, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Enterprise-grade Salesforce file migration and backup tool built with Robot Framework and Python. Supports bulk ContentDocumentId downloa…☆24Updated this week
- Enterprise-grade distributed AI agent framework | Develop → Deploy → Observe | K8s-native | Dynamic DI | Auto-failover | Multi-LLM | Pyth…☆34Updated this week
- GPU-aware inference mesh for large-scale AI serving☆35Sep 25, 2025Updated 9 months ago
- RPC request router and proxy for Starknet, forked from Optimism proxyd.☆12Feb 26, 2024Updated 2 years ago
- A Laravel / Filament starter kit for CMS functionality on websites.☆11Oct 25, 2022Updated 3 years ago
- Example iOS app using the open-source combustion-ios-ble framework.☆11Aug 2, 2023Updated 2 years ago
- Opinionated Elasticsearch query language for Rust☆11Jul 24, 2021Updated 4 years ago
- Intel® AI for Enterprise Inference optimizes AI inference services on Intel hardware using Kubernetes Orchestration. It automates LLM mod…☆44Jul 8, 2026Updated last week
- Groq-powered MAD: The first work to explore Multi-Agent Debate with Large Language Models :D☆12Jul 5, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Source code for 'Professional Sitecore 8 Development' by Phillip Wicklund☆10Mar 27, 2017Updated 9 years ago
- Homebrew tap for openclaw☆43Updated this week
- Mirror of https://gitlab.nic.cz/turris/misc☆12Oct 20, 2021Updated 4 years ago
- ☆37Updated this week
- Evaluate and compare chunking strategies for RAG pipelines☆42May 30, 2026Updated last month
- Just a small collection of packages for Turris OS(OpenWRT) that I build in my free time. Don't expect anything. Not even that it builds..…☆15Jun 10, 2021Updated 5 years ago
- Tree-based, vectorless document RAG framework. Connect any LLM via URL/API key.☆40Apr 7, 2026Updated 3 months ago