A Practitioner handbook for production llm serving.
☆169Aug 5, 2026Updated 2 weeks ago
Alternatives and similar repositories for llm-inference-at-scale
Users that are interested in llm-inference-at-scale are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A physics-grounded, cost-aware optimization loop for vLLM☆64Updated this week
- Dasein managed vector index service - Python SDK☆23May 11, 2026Updated 3 months ago
- Developing K - a language model to generate OPENSCAD code from prompt☆19Dec 3, 2025Updated 8 months ago
- Hands-on Kubernetes troubleshooting lab with 15 production-grade scenarios covering workloads, networking, storage, security, autoscaling…☆23Jul 21, 2026Updated last month
- Become GPU kernel engineer step by step.☆64May 7, 2026Updated 3 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.☆18Aug 14, 2026Updated last week
- ☆17Jul 27, 2026Updated 3 weeks ago
- I developed a fine-tuned retrieval head for RAG that learns to more reliably retrieve relevant passages by transforming the query embeddi…☆17May 21, 2026Updated 3 months ago
- ☆15Mar 30, 2026Updated 4 months ago
- This powerful MCP server bridges the gap between AI assistants and academic research by providing direct access to Semantic Scholar's com…☆16Oct 27, 2025Updated 9 months ago
- One script to replace every outdated tool that Apple ships on macOS with modern versions via Homebrew☆16Jul 18, 2026Updated last month
- A local-first, self-organizing AI RAG GRAPH knowledge system that reads, links and reasons over your documents — offline by default, clou…☆29Apr 22, 2026Updated 4 months ago
- ☆25Dec 15, 2025Updated 8 months ago
- The code implementation of HyGRAG, accepted by WWW'26.☆15May 31, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated 2 months ago
- GPU Engineering for AI Systems☆590May 17, 2026Updated 3 months ago
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆76Jul 6, 2026Updated last month
- Open-source MCP server — progressive tool discovery, code execution, intelligent routing & token optimization across 50+ tools☆15Jun 25, 2026Updated last month
- Programmable automated machine learning - proof of concept☆15Oct 9, 2024Updated last year
- The rag pipeline for optimizing dynamic data editing.☆23Oct 30, 2025Updated 9 months ago
- Demo of fine-tuning QA models for answering FAQ of cloud providers documentation☆11Jun 20, 2026Updated 2 months ago
- An "LLM wiki" upgraded to a real database — typed entities, graph relations, HTTP API, and a built-in natural-language agent.☆106Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆34Mar 24, 2026Updated 4 months ago
- Open Source client side code of INDRZ indoor wayfinding, mapping, routing system. Mirrored from https://gitlab.com/indrz.☆14Feb 16, 2025Updated last year
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆12Nov 26, 2024Updated last year
- A pi extension that adds a /for-$each prompt loop – and hides it from your LLM! Supports children-in-directory and line-in-files iteratio…☆18Jul 20, 2026Updated last month
- Streaming Retrieval-Augmented Generation (RAG) agent in Go. It consumes real-time data from Kafka topics, processes it in configurable wi…☆27Jun 7, 2025Updated last year
- ☆13Apr 16, 2025Updated last year
- Enterprise-grade distributed AI agent framework | Develop → Deploy → Observe | K8s-native | Dynamic DI | Auto-failover | Multi-LLM | Pyth…☆40Updated this week
- A template created for designers, developers, or anyone else who needs a basic and lightweight blog site.☆13Feb 13, 2022Updated 4 years ago
- Reliable RAG setup that uses Semantic Double Merging Chunking from llamaindex, Qdrant Hybrid Search, colBERT for reranking and Google Gem…☆41Dec 15, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- rudradb-opin-examples is for example implementations of the pip install rudradb-opin☆29Mar 3, 2026Updated 5 months ago
- Opinionated Elasticsearch query language for Rust☆11Jul 24, 2021Updated 5 years ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- Homebrew tap for openclaw☆42Updated this week
- CLI for creating new Filament plugins☆13Jun 22, 2026Updated 2 months ago
- Tree-based, vectorless document RAG framework. Connect any LLM via URL/API key.☆42Apr 7, 2026Updated 4 months ago
- This project is a versatile and powerful search tool that leverages state-of-the-art natural language processing models to provide releva…☆12Apr 3, 2023Updated 3 years ago